Code reconstruction method, system and equipment based on function identification and medium
By building a knowledge graph and using a natural language processing model to generate code modification instructions, combined with a graph neural network optimization solution, the efficiency and consistency issues of existing code refactoring tools are solved, and efficient and intelligent code refactoring and optimization are achieved.
Patent Information
- Application Number
- CN202510757238.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-30
AI Technical Summary
Existing code refactoring tools lack efficiency, consistency, and intelligence, making it difficult to automatically identify and modify code implementations, resulting in inconsistencies and quality issues in the code base.
By building a knowledge graph that includes API call chains and data flow paths, combining natural language processing models to analyze user needs, generating a code modification instruction set with weighted annotations, and using graph neural networks to match historical best practice patterns, we generate optimization solutions that retain the original design style and perform static detection and sandbox verification.
It achieves efficient and intelligent code refactoring, improves the consistency and quality of the code base, reduces the risk of human error, and ensures the security and maintainability of the code.
Smart Images

Figure CN120723296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of code optimization, and in particular to a code reconstruction method, system, device and medium based on function identification. Background Art
[0002] In modern software development, large projects are often collaboratively maintained by multiple developers. Due to differences in technical backgrounds, programming habits, and understanding, multiple similar yet distinct code implementations are common for the same functionality. This phenomenon is widespread across codebases, leading to inconsistencies in coding style and implementation. When a specific feature needs to be modified, optimized, or refactored, developers must expend considerable time and effort searching and modifying all relevant code implementations. This manual, file-by-file modification process is not only inefficient but also prone to errors introduced through oversight, compromising code quality and stability.
[0003] Currently, there are some code processing tools on the market, such as code refactoring tools, code checking tools, and static analysis tools. These tools can help developers identify and modify problems in the code to a certain extent, but they have obvious limitations.
[0004] First, in terms of efficiency, most of these tools rely on predefined rules and templates. They are unable to automatically identify and modify all relevant code implementations based on the actual code context and business logic, requiring developers to perform extensive manual intervention and operations, consuming a significant amount of time and effort.
[0005] Secondly, regarding consistency, existing tools lack a global perspective and automation capabilities, making it difficult to ensure consistent coding styles and implementations across the entire codebase. Due to tool limitations, code sections maintained by different developers may not be effectively unified, resulting in numerous inconsistent implementations within the codebase, increasing the difficulty of subsequent maintenance and development.
[0006] Furthermore, existing tools lack intelligence. They rely primarily on simple rule matching and template replacement, lacking a deep understanding of code semantics and context, making it difficult to intelligently modify and optimize code based on complex business needs and logic.
[0007] Finally, in terms of user experience, these tools are often complex to operate, requiring developers to possess certain expertise and skills to master them. Furthermore, the output of these tools may not be intuitive or clear enough to meet the actual needs of developers during development. Summary of the Invention
[0008] The purpose of the present invention is to provide a code reconstruction method, system, device and medium based on function identification, which improves the efficiency and quality of code reconstruction and optimization, and provides software developers with a more convenient, efficient and intelligent code processing tool to solve at least one of the above-mentioned prior art problems.
[0009] In a first aspect, the present invention provides a code refactoring method based on function identification, the method specifically comprising: Scan the project file topology and build a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features; The pre-trained natural language processing model parses the user-entered code function description, combines it with the knowledge graph for context enhancement, and generates a weighted annotated code modification instruction set. Based on the code modification instruction set, combined with data flow sensitivity marking and performance profiling analysis, a three-dimensional quantitative evaluation of the target code in terms of maintainability, security, and execution efficiency is performed to obtain the evaluation results; Based on the evaluation results, a graph neural network is used to match historical best practice patterns and generate an optimization solution that retains the original design style. Static testing of the structural compliance of the optimization scheme's abstract syntax tree is performed. Test cases are injected into the sandbox environment to verify the privacy leakage risks of the optimization scheme during runtime and obtain verification results. By comparing the abstract syntax tree differences between the original code and the verified optimized solution, only functional logic nodes are replaced and legal code style features are retained.
[0010] In a second aspect, the present invention provides a code reconstruction system based on function identification, the system specifically comprising: The structure analysis module is used to scan the project file topology and build a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features; Function recognition module, which uses a pre-trained natural language processing model to parse the user-entered code function description, combines it with the knowledge graph for context enhancement, and generates a weighted annotated code modification instruction set; The quantitative evaluation module is used to perform a three-dimensional quantitative evaluation of the target code in terms of maintainability, security, and execution efficiency based on code modification instruction sets, combined with data flow sensitivity tagging and performance profiling analysis, to obtain evaluation results. The code optimization module is used to match historical best practice patterns based on evaluation results through graph neural networks to generate optimization solutions that retain the original design style; The code verification module is used to check the structural compliance of the abstract syntax tree of the optimization solution at the static level, and inject test cases into the sandbox environment to verify the privacy leakage risk of the optimization solution during operation and obtain verification results; The code replacement module is used to replace only functional logic nodes and retain legal code style features by comparing the abstract syntax tree differences between the original code and the verified optimized solution.
[0011] In a third aspect, the present invention provides a computer device comprising: a memory and a processor and a computer program stored in the memory, wherein when the computer program is executed on the processor, the function identification-based code reconstruction method as described in any one of the above methods is implemented.
[0012] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the code reconstruction method based on function identification as described in any one of the above methods is implemented.
[0013] Compared with the prior art, the present invention has at least one of the following technical effects: 1. The present invention improves the efficiency and quality of code reconstruction and optimization, and provides software developers with a more convenient, efficient and intelligent code processing tool.
[0014] 2. This invention realizes global code reconstruction based on function identification by constructing a knowledge graph including API call chains and data flow paths, and combining it with a natural language processing model to analyze user needs, thereby improving the efficiency and accuracy of code reconstruction while retaining the original design style of the code.
[0015] 3. The present invention analyzes the project file topology, establishes a module hierarchical index, and establishes a four-dimensional relationship network between core entities, thereby enhancing the integrity and accuracy of the knowledge graph and providing a solid foundation for the subsequent generation of code modification instructions.
[0016] 4. The present invention uses a pre-trained natural language processing model to parse user input, and combines it with a knowledge graph for context enhancement to generate a code modification instruction set with weight annotations, thereby realizing intelligent and automated code modification instruction generation and improving the intelligence level of code reconstruction.
[0017] 5. The present invention provides a comprehensive evaluation basis for code refactoring by performing a three-dimensional quantitative evaluation of the target code in terms of maintainability, security, and execution efficiency, which helps developers make more reasonable refactoring decisions and improves code quality.
[0018] 6. This invention matches historical best practice patterns through graph neural networks and generates optimization solutions that retain the original design style, achieving intelligent optimization based on historical experience while maintaining code consistency and maintainability.
[0019] 7. The present invention detects the structural compliance of the optimization scheme at the static level and verifies the privacy leakage risk during runtime in a sandbox environment, ensuring the security and compliance of the optimization scheme and reducing the risk during the reconstruction process.
[0020] 8. By comparing the differences in the abstract syntax trees between the original code and the optimized solution, the present invention only replaces functional logic nodes and retains legal code style features, thereby achieving precise functional modification and style retention, and improving the accuracy of code reconstruction and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a flowchart of a code reconstruction method based on function identification provided by one embodiment of the present invention; Figure 2 This is a schematic structural diagram of a code reconstruction system based on function identification provided by one embodiment of the present invention; Figure 3 It is a structural diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0023] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0024] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0025] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0026] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0027] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0028] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0029] In the embodiments of the present application, the execution subject of the process includes a terminal device, which includes but is not limited to: a server, a computer, a smart phone, a tablet computer, and other devices capable of executing the method disclosed in the present application. Figure 1 A flowchart of a code refactoring method based on function identification disclosed in an embodiment of the present invention is shown, and is described in detail as follows: S101, scans the project file topology structure, and builds a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features.
[0030] In this embodiment, the root directory of the project to be analyzed is specified by user input or configuration file. The root directory is the starting point of the entire project code, and all subsequent file scanning and analysis will be carried out based on this. Starting from the project root directory, all files and subdirectories in the file system are traversed recursively. During the traversal process, the files are preliminarily classified according to their extensions. For example, files such as .java, .py, and .js are identified as source code files, and files such as .xml and .json are identified as configuration files. Different processing strategies are used for different types of files. For each scanned file, its basic information such as the complete path, file name, and file type is recorded. This information will be used to subsequently construct the association relationship and knowledge graph between files.
[0031] Each source code file is parsed using the corresponding parser. For example, Java files can be parsed using the Java compiler API or a third-party parsing library such as ANTLR; Python files can be parsed using Python's built-in ast module. During the parsing of the source code file, all import statements are extracted. Import statements reflect the dependencies between files. For example, a Java file uses import statements to import other classes, and a Python file uses import statements to import other modules.
[0032] Based on the extracted import statements, the association relationship between files is constructed. Each file is regarded as a node in the graph, and the import relationship between files is regarded as a directed edge in the graph. In this way, the file association graph of the entire project is gradually constructed, which can intuitively show the dependency relationship between project files. By analyzing the project files, the technology stack used by the project is identified. For each identified technology stack, its related feature information is extracted. For example, for a SpringBoot project, the SpringBoot version used, the dependent Spring components (such as SpringMVC, SpringData, etc.) and the project configuration method (such as annotation-based configuration, XML-based configuration, etc.) can be extracted; for a Django project, the Django version used, the application configuration information and the database connection settings can be extracted. These technology stack features will provide important contextual information for the subsequent construction of the knowledge graph.
[0033] While parsing the source code, we further analyze the API calls within the code. The form and method of API calls may vary depending on the technology stack. For example, in Java, an API call may appear as a method call or class instantiation; in Python, an API call may appear as a function call or module method call. By analyzing the grammatical structure and semantic information of the code, we identify all API calls.
[0034] Based on the identified API calls, establish call relationships between APIs. Consider each API as a node in a graph, and the call relationships between APIs as directed edges. This way, you can build a chain of API calls within your project, which clearly demonstrates the order and dependencies between API calls.
[0035] Perform data flow analysis on the source code to determine the data flow path within the program. Data flow analysis can be implemented using static analysis techniques. For example, by analyzing the definition, use, and transfer of variables, data can be traced through function calls, conditional statements, and loop statements. Based on the results of data flow analysis, data flow relationships are established. Data sources (such as user input and file reading) and data sinks (such as variable assignments and function return values) are considered nodes in a graph, and the flow of data between nodes is considered directed edges in the graph. In this way, a data flow path is constructed within the project, which can intuitively demonstrate the direction of data flow and processing within the program.
[0036] Integrate the previously constructed file association graph, API call graph, and data flow path graph to build a unified knowledge graph. During the integration process, ensure that the nodes and edges in the different graphs are correctly associated. For example, file nodes can be associated with the API nodes defined within them, and API nodes can be associated with related nodes in the data flow path.
[0037] Add rich attribute information to each node and edge in the knowledge graph. For example, for file nodes, you can add attributes such as the file's creation time, modification time, and author; for API nodes, you can add attributes such as the API's function description, parameter information, and return value information; for edges, you can add attributes such as edge type (such as import relationship, call relationship, data flow relationship, etc.) and weight (such as call frequency, data transfer volume, etc.). This attribute information will provide more comprehensive contextual support for subsequent code analysis and refactoring.
[0038] In this example, we were able to successfully scan the project file topology and, based on code import relationships and technology stack features, construct a knowledge graph containing API call chains and data flow paths. This knowledge graph clearly displays the overall structure and logical relationships of the project code, providing strong support for subsequent code refactoring based on function identification. Developers can use this knowledge graph to quickly understand the code's functions and dependencies, improving the efficiency and quality of code analysis and refactoring.
[0039] S102: parse the user-entered code function description through a pre-trained natural language processing model, and perform context enhancement in combination with the knowledge graph to generate a code modification instruction set with weight annotations.
[0040] In this embodiment, an input interface is provided for developers, allowing them to enter a code function description through a text box or file upload. The code function description can be in natural language and detail the code function to be modified, such as "Change the user login verification logic from database-based to third-party authentication service-based verification, while keeping the original error message unchanged." The code function description entered by the user is initially preprocessed, including removing unnecessary spaces, line breaks, special characters, etc., and converting the text into a unified format to facilitate subsequent natural language processing model parsing.
[0041] The preprocessed code description is fed into a pretrained natural language processing model. The model conducts a deep analysis of the input text, extracting key information such as the code module requiring modification, the specific operations being modified, and the business logic involved. For example, the model can identify key concepts such as "user login verification logic," "third-party authentication service," and "error message" in the user's description. Based on the model's analysis, preliminary code modification instructions are generated. These instructions describe the general direction and content of the code modification in natural language, such as "Replace the database verification logic in the user login verification module with the third-party authentication service call logic, and ensure that the error message is consistent with the original logic."
[0042] The preliminarily generated code modification instructions are associated with the previously constructed knowledge graph for query. The knowledge graph contains information such as the topological structure of the project file, the API call chain, the data flow path, etc. By querying the knowledge graph, contextual information related to the code modification instructions can be obtained. The queried knowledge graph contextual information is merged with the preliminarily generated code modification instructions. By analyzing the contextual information, the specific scope and details of the code modification are further clarified. For example, if the knowledge graph shows that the user login verification module also has data interaction with other modules, then these interactions need to be considered in the generated code modification instructions to ensure that the modified code does not destroy the original data flow and business logic.
[0043] The preliminary instructions are enhanced based on contextual information to include more detailed and accurate code modification information. For example, the instructions clearly indicate the code file path to be modified, the specific code line number, and the code snippet to be replaced or added.
[0044] A weight index is determined for each instruction in the code modification instruction set. The weight index can be set based on factors such as the importance of the instruction, the scope of influence, and the difficulty of execution. Based on the determined weight index, the enhanced code modification instructions are weighted. The weight calculation can adopt a comprehensive scoring method, taking into account multiple factors such as the code modules involved in the instruction, the impact on system performance, and the degree of correlation with other modules. After the calculation is completed, the weight is marked on each instruction to form a code modification instruction set with weight marking. The code modification instruction set with weight marking is sorted and sorted, and the instructions are arranged in order from high to low weight. The sorted instruction set is output to the subsequent code reconstruction module so that the code modification operation can be performed according to the priority of the instruction.
[0045] In this embodiment, a pre-trained natural language processing model can parse the user-entered code function description, and combined with the knowledge graph for contextual enhancement, it generates a weighted and annotated code modification instruction set. This instruction set not only accurately reflects the developer's code modification intentions but also includes rich contextual information and weight annotations, providing precise and efficient guidance for subsequent code refactoring, improving the quality and efficiency of code refactoring.
[0046] S103, based on the code modification instruction set, combined with data flow sensitivity marking and performance profiling analysis, a three-dimensional quantitative evaluation of the target code in terms of maintainability, security, and execution efficiency is performed to obtain the evaluation results.
[0047] In this embodiment, a previously generated set of weighted code modification instructions is analyzed in detail. The specific code modules, functions, variables, and other information involved in each instruction are clearly identified. For example, if the instruction is "Modify the password encryption logic in the user login verification module," the target code is determined to include all code files related to the user login verification module and functions involving password encryption logic. Based on the analysis results, the specific scope of the target code is circled in the project code library. Using identifiers such as file paths, function names, and class definitions, the target code to be evaluated is precisely determined.
[0048] Extract the data flow path information related to the target code from the knowledge graph. Clarify the input source, processing process, and output target of the data in the target code. Sensitivity-label the data flow in the target code based on the nature and importance of the data. Categorize the data into different sensitivity levels, such as high sensitivity (such as user passwords, personal identity information), medium sensitivity (such as user preferences), and low sensitivity (such as page visit statistics). Consider the confidentiality, integrity, and availability requirements of the data when labeling. Associate the labeled data flow sensitivity information with specific code segments in the target code. Determine which code segments process high-sensitivity data and which process medium-sensitivity or low-sensitivity data. This will help focus on analyzing data processing codes of different sensitivities in subsequent evaluations.
[0049] Collect performance data from the target code during its historical execution. This data can be obtained through performance monitoring tools, logging, and other methods. Performance data includes metrics such as code execution time, memory usage, and CPU utilization. Build a performance profile of the target code based on the collected performance data. Performance profiles intuitively display the code's performance in different scenarios. Charts (such as line charts and bar charts) or visualization tools can be used to present the changing trends and distribution of performance metrics. By analyzing the performance profiles, performance bottlenecks in the target code can be identified. Identify code segments with excessive execution time, excessive memory usage, or abnormal CPU usage. For example, record the average execution time and maximum memory usage of the user login verification module under different load conditions, plot a curve showing how the user login verification module's execution time changes as the number of users increases, and create a bar chart showing memory usage under different operations. If the password encryption function's execution time increases significantly during peak user login periods, slowing down the entire login process, this function may be a performance bottleneck.
[0050] Identify metrics for evaluating the maintainability of the target code, such as code complexity (including cyclomatic complexity and class coupling), code comment coverage, and code duplication. Code complexity reflects the readability and understandability of the code; highly complex code is often difficult to maintain. Code comment coverage reflects the explainability of the code; good comments help other developers understand the code logic. Code duplication reflects the degree of code redundancy; duplicate code increases maintenance costs. Use appropriate tools or algorithms to calculate the maintainability index value of the target code. Based on the calculated maintainability index value, assess the maintainability of the target code. Maintainability can be categorized into five levels: excellent, good, fair, poor, and poor. For example, if code complexity is low, comment coverage is high, and duplication is low, the maintainability is assessed as excellent. Conversely, if code complexity is high, comments are insufficient, and there is a large amount of duplicate code, the maintainability is assessed as poor.
[0051] Combined with data flow sensitivity tags, identify possible security risk points in the target code. Focus on code segments that process highly sensitive data and check for security vulnerabilities such as SQL injection, cross-site scripting (XSS), and plaintext password storage. Use security scanning tools to scan the target code to detect known security vulnerabilities and potential security issues. Security scanning tools can conduct a comprehensive inspection of the code based on predefined security rules and vulnerability libraries. Based on the identified security risk points and security scanning results, conduct a security level assessment of the target code. Security is divided into three levels: high, medium, and low. If there are serious security vulnerabilities in the code and no effective protective measures are taken, the security assessment is low; if the code basically complies with security specifications and only has a few minor security issues, the security assessment is medium; if the code performs well in terms of security and has taken comprehensive security protection measures, the security assessment is high.
[0052] Combined with the performance profile, analyze the execution efficiency of the target code in different scenarios. Focus on the execution time and resource usage of the performance bottleneck code segment. Compare the execution efficiency of the target code with the industry benchmark. The industry benchmark can be the average performance index of the same type of project or code with similar functions under the same hardware environment and load conditions. Through comparison, determine the advantages and disadvantages of the target code in terms of execution efficiency. For example, for the password encryption function that shows an excessively long execution time in the performance profile, analyze the reasons for its low execution efficiency, such as improper algorithm selection, too many loops, etc. If the execution time of the user login verification module in the target code is significantly higher than the industry benchmark, it means that its execution efficiency needs to be improved.
[0053] Based on the comparison results, the target code's execution efficiency is assessed. Execution efficiency is categorized into three levels: high, average, and low. If the target code's execution efficiency is better than the industry benchmark or within a reasonable range, it is assessed as high. If it is similar to the industry benchmark but still has room for improvement, it is assessed as average. If the execution efficiency is significantly lower than the industry benchmark, it is assessed as low.
[0054] The evaluation results of the three dimensions of maintainability, security, and execution efficiency are integrated to form a comprehensive evaluation report, which includes the evaluation indicator value, evaluation level, and detailed analysis description of each dimension.
[0055] This embodiment, based on the code modification instruction set, combines data flow sensitivity tagging with performance profiling analysis to perform a three-dimensional quantitative assessment of the target code's maintainability, security, and execution efficiency, obtaining comprehensive and accurate assessment results. This assessment provides an important basis for subsequent code refactoring, allowing developers to develop targeted optimization plans based on the assessment results to improve code quality, security, and performance.
[0056] S104, based on the evaluation results, matches historical best practice patterns through graph neural networks to generate an optimization solution that retains the original design style.
[0057] In this embodiment, multiple completed and well-running historical code projects are collected from channels such as the company's internal code repository and open source project platforms. These projects cover the same or similar technology stacks, business areas, and functional modules as the current target code project to ensure that the historical best practice patterns have reference value. The collected historical code projects are deeply analyzed to extract the code patterns therein. The code pattern can be a specific algorithm implementation, design pattern application, code structure organization method, etc. The extracted code patterns are sorted and classified to build a historical best practice pattern library. Detailed description information is added to each pattern, including the application scenarios, advantages, applicable conditions, etc. of the pattern. At the same time, the pattern is represented in the form of a graph structure, where the nodes in the graph represent elements such as classes, functions, variables, etc. in the code, and the edges represent the calling relationships, dependencies, etc. between them.
[0058] Extract key indicator features from the three-dimensional quantitative evaluation results of the target code's maintainability, security, and execution efficiency. Maintainability indicator features may include code complexity, comment coverage, code duplication rate, etc.; security indicator features may include the number of security vulnerabilities, the completeness of security protection measures, etc.; execution efficiency indicator features may include code execution time, resource usage, etc. Quantify the extracted evaluation indicator features, and convert non-numerical features (such as the completeness of security protection measures) into numerical feature values. Quantification can be performed using methods such as grading and scoring. Combine the quantified evaluation indicator features into a feature vector as a digital representation of the target code evaluation results. Each dimension of the feature vector corresponds to a feature value of an evaluation indicator feature.
[0059] The pattern graph structures in the historical best practice pattern library are associated with the corresponding evaluation indicator feature vectors to construct a training dataset. Each training sample contains a pattern graph structure and its corresponding evaluation indicator feature vector, as well as the comprehensive performance score (e.g., excellent, good, fair, poor) achieved after the pattern was applied in a real project. A graph neural network architecture suitable for pattern matching tasks is designed. The graph neural network should be able to process graph-structured data, extract feature information from the graph, and fuse it with the evaluation indicator feature vector. Common graph neural network architectures include graph convolutional networks (GCNs) and graph attention networks (GATs). The graph neural network model is trained using the training dataset. During training, the model parameters are adjusted to enable the model to accurately predict the comprehensive performance score after the historical best practice pattern is applied in real projects. Methods such as cross-validation are used to evaluate and optimize the model to prevent overfitting.
[0060] The feature vector of the target code evaluation result is input into the trained graph neural network model. Based on the information in the feature vector, the model filters and matches patterns in the historical best practice pattern library. The graph neural network model evaluates each historical best practice pattern and calculates its match with the target code evaluation result. This match is calculated based on the pattern's alignment with the target code requirements in terms of maintainability, security, and execution efficiency. Based on the match results, several historical best practice patterns with the highest match to the target code evaluation result are selected. These patterns serve as the basis for candidate optimization solutions.
[0061] Analyze the original design style of the target code. The design style includes aspects such as code naming conventions, code structure organization, and design pattern application. Integrate the best-matching historical best practice pattern with the original design style of the target code. During the fusion process, retain the reasonable parts of the original design style, and apply the optimization points in the best-matching pattern to the target code. For example, if the class naming in the target code uses camelCase naming, and the class naming in the best-matching pattern uses underscore naming, then when generating the optimization plan, maintain the camelCase naming of the target code, and only apply the algorithm logic and code structure optimization parts in the best-matching pattern to the target code.
[0062] Based on the fusion results, an optimization plan is generated that retains the original design style. The optimization plan should describe in detail the specific modifications to the target code, including code additions, deletions, and modifications.
[0063] Perform a preliminary evaluation of the generated optimization solution to check whether it meets the target code's requirements for maintainability, security, and execution efficiency, while also preserving the original design style. This preliminary evaluation can be performed manually or using simple code review tools. If the preliminary evaluation reveals issues with the optimization solution, such as not fully meeting the improvement requirements or undermining the original design style, adjust the optimization solution based on the feedback. Repeat these steps until a satisfactory optimization solution is generated.
[0064] In this embodiment, based on the evaluation results, a graph neural network can be used to match historical best practice patterns and generate an optimization solution that retains the original design style. This optimization solution not only effectively improves the maintainability, security, and execution efficiency of the target code, but also aligns with the overall project code style, reducing the difficulty of subsequent maintenance and development, and improving the quality and efficiency of code refactoring.
[0065] S105, detecting the structural compliance of the abstract syntax tree of the optimization scheme at the static level, and injecting test cases into the sandbox environment to verify the privacy leakage risk during the operation of the optimization scheme, and obtaining the verification results.
[0066] In this embodiment, an appropriate abstract syntax tree (AST) generation tool is selected based on the programming language used in the project. For example, for Python projects, the ast module can be used; for Java projects, the AST generation function provided by Eclipse JDT (Java Development Tools) can be used. These tools can parse source code into a tree structure, where each node represents a syntactic element in the code, such as a variable declaration, function call, or control statement.
[0067] Use the selected AST generation tool to parse the generated optimized code. The tool reads the code line by line, identifies the syntax elements, and constructs an abstract syntax tree according to the syntax rules. The generated abstract syntax tree is output in a suitable format for subsequent structural compliance checking. The AST can be saved to a file or directly constructed into a data structure in memory that can be accessed by the program.
[0068] Define a set of structural compliance rules for the abstract syntax tree based on the project's coding standards and best practices. Develop an abstract syntax tree structure detection tool based on the defined structural compliance rules. This tool can traverse the nodes of the abstract syntax tree and check whether each node complies with the corresponding rules. Use the developed structural detection tool to detect the abstract syntax tree of the optimization solution. The tool checks each node of the AST according to predefined rules and records all violations. Based on the detection results, a detailed structural compliance report is generated. The report contains all detected violation information, including the location of the violating node (such as file path, line number), the specific rule of the violation, and a description of the violation.
[0069] Select a suitable sandbox platform based on the project's operating environment and requirements. The sandbox platform should be able to provide an isolated operating environment to prevent the optimization solution from affecting the actual system during the testing process. Common sandbox platforms include Docker containers, virtual machines, etc. Configure an environment similar to the actual operating environment of the project on the selected sandbox platform. This includes installing the necessary software packages, dependent libraries, and runtime environment. To ensure the security of the sandbox environment, set appropriate security policies. For example, limit network access rights within the sandbox to prevent the optimization solution from accessing external sensitive resources during the testing process; set resource usage limits, such as CPU usage, memory usage, etc., to prevent programs within the sandbox from occupying too many system resources.
[0070] Design a series of test cases to verify privacy leakage risks based on the project's business logic and potential privacy information. Inject the designed privacy leakage test cases into the established sandbox environment. Test cases can be automatically executed by writing test scripts or using a test framework (such as Python's unittest, pytest, etc.). Execute the injected test cases in the sandbox environment and monitor the running status of the optimization solution in real time. Monitoring content includes program output information, logging, system resource usage, etc. Based on the monitoring information and program output results during test execution, analyze whether the optimization solution has privacy leakage risks. If the program is found to exhibit abnormal behavior when processing certain test cases, such as outputting sensitive information or generating abnormal logs, there may be a privacy leakage issue.
[0071] Integrate the results of statically checking the structural compliance of the abstract syntax tree and the results of verifying privacy leakage risks in the sandbox environment. Generate a comprehensive verification report that includes the violation information found in the structural compliance test and the potential risks found in the privacy leakage test. Based on the integrated verification results, evaluate the feasibility of the optimization plan. If no issues are found in the structural compliance test and the privacy leakage test, the optimization plan is considered to meet the requirements in terms of structure and security, and subsequent development and deployment work can continue. If there are violation information or potential risks, the optimization plan needs to be modified and improved according to the specific situation.
[0072] This embodiment effectively checks the structural compliance of the optimization scheme's abstract syntax tree at a static level, identifying and resolving code structure issues in advance. Furthermore, it accurately verifies the privacy leakage risks of the optimization scheme during runtime in a sandbox environment, ensuring code security. The resulting verification results provide developers with comprehensive and accurate feedback, helping to improve the quality and reliability of code refactoring and reduce subsequent development and maintenance costs.
[0073] S106, by comparing the abstract syntax tree differences between the original code and the verified optimization solution, only functional logic nodes are replaced and legal code style features are retained.
[0074] In this example, an appropriate Abstract Syntax Tree (AST) generation tool is selected based on the project's programming language. The selected AST generation tool is used to parse the original code and the verified optimized solution code. The tool reads the code line by line, identifies the syntax elements, and constructs an Abstract Syntax Tree (AST) according to the syntax rules. The generated ASTs for the original code and the optimized solution are output in an appropriate format for subsequent comparative analysis.
[0075] Based on the grammatical structure and semantics of the code, define a set of node matching rules. These rules are used to determine which nodes in the AST of the original code and the optimized solution correspond, that is, which nodes have the same or similar functions. Based on the defined node matching rules, develop an abstract syntax tree comparison tool. This tool can traverse the AST of the original code and the optimized solution, compare them node by node, and find the differences between them. Use the developed AST comparison tool to compare the abstract syntax trees of the original code and the optimized solution. The tool will check each node of the AST according to predefined rules and record all mismatched nodes. Based on the comparison results, generate a detailed difference report. The report should contain all detected difference information, including the location of the difference node (such as file path, line number), the type of difference (such as new node, deleted node, modified node) and a description of the difference.
[0076] Based on the project's business requirements and coding standards, define the characteristics of functional logic nodes and coding style characteristic nodes. Functional logic nodes are code nodes that implement specific business functions, such as function calls, conditionals, and loop statements. Code style characteristic nodes are nodes related to coding style, such as indentation, spacing, and comments. Analyze the difference report and identify each difference node as a functional logic node or a coding style characteristic node based on the defined characteristics.
[0077] Based on the abstract syntax tree of the original code, an AST framework for the target code is constructed. This framework retains all legal code style feature nodes in the original code, such as indentation, spaces, and comments. Based on the difference report, the functional logic nodes in the optimization solution are replaced with the corresponding locations in the AST framework of the target code. During the replacement process, the functional logic nodes are ensured to be correctly connected with other nodes in the AST framework of the target code to maintain the syntactic correctness of the code. The AST framework of the target code is converted back to source code using an AST generation tool. The tool generates the corresponding code strings based on the node information in the AST and the grammatical rules of the programming language.
[0078] In this embodiment, the abstract syntax tree differences between the original code and the verified optimization solution are effectively compared, replacing only the functional logic nodes while retaining legal coding style features. This allows code functionality to be optimized while maintaining the consistency of the project's original coding style, reducing the difficulty of subsequent maintenance. Furthermore, by verifying the target code, its quality and stability are ensured, improving the success rate and efficiency of code refactoring.
[0079] In some embodiments, in the above step S101, scanning the project file topology structure and constructing a knowledge graph including API call chains and data flow paths based on code import relationships and technology stack features specifically includes: Analyze the project file topology and establish a module hierarchy index based on file path dependencies; Identify core entities from the source code based on the module level index, wherein the core entities include API definitions, data objects, and external dependencies; Based on the code call stack and data flow analysis, a four-dimensional relationship network including call, inheritance, transmission, and dependency is established between core entities; According to the technology stack declaration in the project configuration file, framework constraint rules and best practice templates are injected into the four-dimensional relationship network to obtain an enhanced four-dimensional relationship network; The enhanced four-dimensional relationship network is converted into an attribute graph structure, wherein the nodes of the attribute graph structure are used to carry code location identifiers, and the edges of the attribute graph structure are used to mark relationship types and call frequencies.
[0080] In this embodiment, a file scanning tool is developed that can traverse the project directory and collect the path information of all code files. The tool can recursively scan the subdirectories under the project directory to ensure that no files are missed. For each collected code file, its path dependency with other files is analyzed. By viewing the import statements in the file, the dependency hierarchy between files is determined. According to the file path dependency, related files are organized into modules, and a module hierarchy index is constructed. A tree structure can be used to represent the module hierarchy, where the root node represents the entire project, the child nodes represent different modules, and the leaf nodes represent specific code files.
[0081] Clarify the types of core entities, including API definitions, data objects, and external dependencies. API definitions refer to the interfaces provided by the project, such as functions and methods; data objects refer to the data structures used in the project, such as classes and structures; and external dependencies refer to the external libraries or frameworks that the project relies on.
[0082] Develop a core entity recognition tool based on the module-level index. This tool can parse syntactic elements in the source code to identify parts that meet the core entity definition. For example, for API definitions, the tool can search for function declarations and method definitions in the source code to determine whether they are external interfaces. For data objects, the tool can search for class definitions and structure definitions. For external dependencies, the tool can analyze import statements and identify dependent external libraries or frameworks.
[0083] Use the developed core entity recognition tool to scan the project source code and identify all core entities. The tool will record the identified core entity information, including the core entity name, file path, module, etc.
[0084] Develop a code analysis tool to analyze the code's call stack and data flow. This tool can analyze function call statements, variable assignment statements, and other statements in the source code to determine the calling relationships, inheritance relationships, data transfer relationships, and dependency relationships between core entities. For example, by analyzing function call statements, you can determine the calling relationships between APIs; by analyzing class inheritance statements, you can determine the inheritance relationships between classes; by analyzing variable assignments and transfers, you can determine the data transfer relationships between data objects; and by analyzing import statements and class usage, you can determine the dependency relationships between core entities.
[0085] Based on the results of the code analysis tool, a four-dimensional relationship network consisting of call, inheritance, transfer, and dependency is established between core entities. This relationship network can be represented using a graph structure, where nodes represent core entities and edges represent relationships between core entities. Edge types include call edges, inheritance edges, transfer edges, and dependency edges.
[0086] Develop a configuration file parsing tool to parse the technology stack declarations in project configuration files. This tool can extract information about the frameworks and libraries used in the project, as well as their version numbers. Based on the parsed technology stack information, collect corresponding framework constraints and best practice templates. Framework constraints refer to the code writing standards and restrictions specified by the framework, such as annotation usage rules in the Spring framework and entity mapping rules in the Hibernate framework. Best practice templates refer to the recommended code structure and implementation methods when developing within the framework.
[0087] Develop a rule injection tool to inject the collected framework constraint rules and best practice templates into the four-dimensional relationship network. The corresponding rules and template information can be added to the nodes or edges of the relationship network.
[0088] Define the required attribute information for nodes and edges in the property graph structure. Nodes carry code location identifiers, such as file paths and line numbers, allowing developers to quickly locate specific locations within the code. Edges are used to annotate relationship types and call frequencies. Relationship types include call, inheritance, transfer, and dependency. Call frequencies can be determined by analyzing code call records or log information. Based on the enhanced four-dimensional relationship network, develop a property graph structure conversion tool. This tool converts the node and edge information in the relationship network into the required format for the property graph structure. For example, nodes in the relationship network can be converted into nodes in the property graph structure and assigned code location identifier attributes; edges in the relationship network can be converted into edges in the property graph structure and assigned relationship type and call frequency attributes. Use the developed property graph structure conversion tool to process the enhanced four-dimensional relationship network to generate a property graph structure. The generated property graph structure can be saved as a file, perhaps using a graph database storage format (such as Neo4j's Cypher statement format) for subsequent querying and analysis.
[0089] This embodiment effectively scans the project file topology and constructs a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features. This knowledge graph is presented as a property graph structure, with nodes carrying code location identifiers and edges annotated with relationship types and call frequencies. This facilitates developers to quickly understand and analyze project code structures, providing strong support for code optimization, refactoring, and development.
[0090] In some embodiments, in step S102 above, the code function description input by the user is parsed by a pre-trained natural language processing model, and contextually enhanced in combination with a knowledge graph to generate a code modification instruction set with weighted annotations, specifically including: A Transformer-based bidirectional encoder is used to parse the functional description input by the user, and a code function semantic vector is generated through the semantic encoding layer; Based on the code function semantic vector, match related entities and relationship subgraphs from the knowledge graph; Using the retrieved subgraph structure, we perform graph attention weighting on the code function semantic vector to generate a context-enhanced semantic vector. Parsing an atomic operation sequence according to the context-enhanced semantic vector, where each atomic operation in the atomic operation sequence includes a target code location and a modification type; Based on the entity sensitivity and call frequency in the knowledge graph, the execution priority weights are assigned to the atomic operations to obtain a weighted atomic operation sequence; Encapsulate weighted atomic operation sequences into a machine-executable code modification instruction set.
[0091] In this embodiment, a bidirectional encoder model based on the Transformer architecture is selected, such as the BERT (Bidirectional Encoder Representations from Transformers) series of models. The BERT model uses a bidirectional attention mechanism to simultaneously consider the context on both the left and right sides of a word, thereby more accurately understanding the meaning of the word. An input processing module is designed to receive the code function description input by the user. The user input is preprocessed, including removing irrelevant characters, unifying capitalization, and performing word segmentation. The preprocessed user input is input into the Transformer-based bidirectional encoder. The model processes the input vocabulary sequence through its internal semantic encoding layer, maps each word to a high-dimensional vector space, and comprehensively considers the semantic information of the entire sequence to ultimately generate a semantic vector representing the code function description. This semantic vector contains the core semantic information of the user input, such as functional requirements, operation objects, etc.
[0092] Develop a knowledge graph query interface that can interact with a pre-built knowledge graph. The knowledge graph stores various entities and relationships related to the project code, such as API definitions, data objects, call relationships, etc. Use the generated code function semantic vector to perform entity matching in the knowledge graph. Relevant entities can be found by calculating the similarity between the semantic vector and the entity vector in the knowledge graph. For example, using the cosine similarity algorithm, calculate the cosine value of the angle between the code function semantic vector and each entity vector in the knowledge graph. Entities with high similarity are entities related to the user-input function description. For matched entities, retrieve their associated relationship subgraph from the knowledge graph. The relationship subgraph contains various relationships between the entity and other entities, such as call relationships, dependency relationships, etc.
[0093] A graph attention mechanism module was developed to analyze the retrieved subgraph structure. The graph attention mechanism assigns different weights to the code function semantic vector based on the importance of different nodes (entities) in the subgraph. For each node in the relational subgraph, the node's weight is calculated based on factors such as its similarity to the code function semantic vector and the node's centrality within the subgraph. Based on the calculated node weights, the code function semantic vectors are weighted and fused to generate a context-enhanced semantic vector. This vector not only incorporates the original semantic information entered by the user but also incorporates the contextual information of related entities in the knowledge graph, making the semantic representation richer and more accurate.
[0094] Clarify the definition and type of atomic operations. Atomic operations are the smallest unit of code modification, such as adding a line of code, deleting a line of code, or modifying a variable value. Based on the context-enhanced semantic vector, develop an operation parsing module. This module analyzes and understands the semantic vector and parses it into a series of atomic operations. For each parsed atomic operation, determine its target code location and modification type. The target code location can be determined by analyzing the code location information in the knowledge graph and the relationship between the atomic operation and the code entity. The modification type is the atomic operation type defined above. For example, if the context-enhanced semantic vector indicates "adding verification code verification to the login function," the operation parsing module might parse out atomic operations such as "adding verification code generation code to the beginning of the login method," "adding verification code parameters to the login method parameters," and "adding verification code verification logic to the login method logic." For the atomic operation "adding verification code generation code to the beginning of the login method," the target code location is the beginning of the login method, and the modification type is code addition.
[0095] In the knowledge graph, sensitivity and call frequency indicators are defined for each entity. Entity sensitivity reflects the importance of the entity in the project and the degree of its impact on system stability. For example, entities related to core business logic have higher sensitivity. Call frequency indicates the number of times an entity is called in the code. Entities with high call frequencies usually have a greater impact on system functions. Based on the sensitivity and call frequency of the entities involved in the atomic operation in the knowledge graph, an execution priority weight is assigned to each atomic operation. For example, if an atomic operation involves an entity with high sensitivity and a high call frequency, then the atomic operation will be given a higher weight, indicating that its execution priority is higher. The calculated execution priority weight is assigned to the corresponding atomic operation to form a weighted atomic operation sequence. This sequence is sorted from high to low according to weight, and developers can execute atomic operations in sequence according to the weight order to ensure that important modifications are performed first.
[0096] Define a machine-executable code modification instruction set format that includes information such as the target code location, modification type, and execution priority weight of the atomic operation. For example, the instruction set can use JSON format, with each atomic operation as a JSON object containing fields such as "position" (target code location), "type" (modification type), and "weight" (execution priority weight).
[0097] Develop an instruction encapsulation module to encapsulate weighted atomic operation sequences according to a predefined instruction set format. This module traverses the atomic operation sequence, converts the information of each atomic operation into an instruction set representation, and combines them into a complete instruction set. The instruction encapsulation module processes the weighted atomic operation sequence to generate the final machine-executable code modification instruction set. This instruction set can be used by automated code modification tools or other related systems to achieve automated code modification.
[0098] In this embodiment, a pre-trained natural language processing model can effectively parse user-entered code function descriptions, combine them with knowledge graphs for contextual enhancement, and generate a weighted, annotated set of code modification instructions. This instruction set provides clear guidance and prioritization for code modifications, improving the efficiency and accuracy of code modifications, reducing the possibility of manual intervention and errors, and contributing to improved software development quality and efficiency.
[0099] In some embodiments, in step S103 above, the target code is subjected to a three-dimensional quantitative evaluation of maintainability, security, and execution efficiency based on the code modification instruction set, combined with data flow sensitivity tagging and performance profile analysis, to obtain an evaluation result, which specifically includes: Extract the code segment to be evaluated based on the target code location and operation type in the code modification instruction set; Track the first sensitive information propagation path of the code segment to be evaluated through a pre-built data flow graph, and calculate the security risk score based on the first sensitive information propagation path using the privacy rule base; Execute the code segment to be evaluated in the sandbox environment, collect resource consumption indicators, establish an execution efficiency profile model, and output the execution efficiency profile; Based on the code complexity index and change history data, the maintainability prediction model is used to analyze the code segment to be evaluated and generate the evolution cost coefficient; The code complexity indicators include cyclomatic complexity, change cost, and dependency entropy. The cyclomatic complexity is calculated by the McCabe index, the change cost is calculated by multiplying the file modification frequency and the impact range, and the dependency entropy is calculated by multiplying the number of external dependencies and the version dispersion. The security risk score, execution efficiency profile, and evolution cost coefficient are input into the evaluation matrix to generate a three-dimensional evaluation result including a weight vector.
[0100] In this embodiment, an instruction set parsing tool is developed to read the code modification instruction set. The instruction set contains information such as the target code location and operation type. For example, the instruction set may clearly indicate the code addition, modification, or deletion operation at a specific line in a file. The parsing tool analyzes the instruction set line by line, extracting the target code location and operation type involved in each instruction. Based on the parsed target code location, the corresponding code segment is located in the project code library.
[0101] A data flow graph is pre-built. This graph documents the flow of data from input to output within the project code, as well as the data transfer relationships between different modules and functions. Static code analysis tools can be used to analyze variable assignments, function calls, and other statements within the code to construct the data flow graph. For the extracted code segment to be evaluated, the first sensitive information propagation path associated with that code segment is tracked within the data flow graph. Sensitive information may include user privacy data, critical system configurations, and so on.
[0102] A privacy rule library is established, containing various sensitive information protection rules and security risk assessment criteria. For example, the rule library stipulates that password data must be encrypted during transmission, otherwise it will pose a high security risk. Based on the traced transmission path of the first sensitive information, the code segment to be evaluated is checked for compliance with security requirements against the rules in the privacy rule library. For paths or operations that do not comply with the rules, a corresponding security risk score is assigned according to the scoring criteria in the rule library. Ultimately, the risk scores of all non-compliance cases are combined to obtain an overall security risk score for the code segment to be evaluated.
[0103] Build a sandbox environment similar to the actual production environment. This environment should have the same hardware configuration, operating system, and runtime library as the production environment. The sandbox environment can isolate the execution of the code segment to be evaluated to avoid affecting the actual system. Execute the code segment to be evaluated in the sandbox environment, and use monitoring tools to collect resource consumption indicators during code execution, such as CPU usage, memory usage, disk I / O, etc. The monitoring tool can record changes in these indicators in real time, for example, collecting data at regular intervals. Based on the collected resource consumption indicators, establish an execution efficiency profile model. This model can display the resource consumption of the code at different execution stages. For example, it can draw a curve chart showing the change of CPU usage over time, and the correspondence between memory usage and code execution steps. The execution efficiency profile model can intuitively understand the execution efficiency bottlenecks of the code, such as a function that takes too long to execute or abnormal memory usage.
[0104] Cyclomatic complexity is calculated using the McCabe index. The McCabe index is a measure of code complexity based on the number of control flow structures (such as conditionals and loops) in the code. For example, a function containing multiple if-else statements and for loops would have a high cyclomatic complexity, indicating that the function's control flow is complex and difficult to understand and maintain.
[0105] The cost of a change is calculated by multiplying the file's modification frequency and its scope of impact. The frequency of a file's modification can be determined by counting the number of times the file has been modified in the codebase, while the scope of impact can be determined by analyzing the file's dependencies on other modules or functions. For example, a file that is frequently modified and affects multiple other modules will have a higher cost of change.
[0106] Dependency entropy is calculated by multiplying the number of external dependencies by their version dispersion. The number of external dependencies refers to the number of external libraries or frameworks used in the code, while version dispersion refers to the number of different versions of these external dependencies. For example, a project using multiple different versions of external libraries will have a high dependency entropy, indicating a complex dependency relationship and greater difficulty in management and maintenance.
[0107] Based on historical code data, a maintainability prediction model is constructed. This model uses machine learning algorithms, such as decision trees and neural networks, to predict the code's evolutionary cost coefficient using input features such as code complexity metrics and change history data. Information such as the code complexity metric of the code segment to be evaluated is fed into the maintainability prediction model. The model uses the trained algorithm to predict the code's evolutionary cost and generates an evolutionary cost coefficient. This coefficient reflects the effort and difficulty required to modify, expand, or maintain the code in the future.
[0108] Design an evaluation matrix that integrates the security risk score, execution efficiency profile, and evolution cost coefficient as three dimensions. The evaluation matrix can be a multidimensional data structure used to comprehensively evaluate the maintainability, security, and execution efficiency of the code. Based on the actual needs and focus of the project, determine the weight vectors of the security risk score, execution efficiency profile, and evolution cost coefficient in the evaluation matrix. Input the security risk score, execution efficiency profile, and evolution cost coefficient into the evaluation matrix, and combine them with the determined weight vectors for comprehensive calculation and evaluation. Finally, a three-dimensional evaluation result containing the weight vector is generated, which intuitively shows the comprehensive performance of the code in terms of maintainability, security, and execution efficiency.
[0109] In this embodiment, based on the code modification instruction set, combined with data flow sensitivity tagging and performance profiling analysis, a three-dimensional quantitative assessment of the target code's maintainability, security, and execution efficiency is performed, yielding accurate assessment results. This assessment helps developers fully understand the code quality, identify existing problems and potential risks, and provide a scientific basis for code optimization, refactoring, and decision-making, thereby improving the quality and efficiency of software development and maintenance.
[0110] In some embodiments, in step S104 above, based on the evaluation results, matching historical best practice patterns through a graph neural network to generate an optimization solution that retains the original design style specifically includes: Extract successful refactoring cases based on the historical version library and convert the abstract syntax tree differences before and after the code changes of the successful refactoring cases into graph structure samples; A graph neural network model is built based on the graph structure sample. The abstract syntax tree of the target code and the evaluation results are input into the graph neural network model for matching. The top K best practice candidate solutions are screened from the historical knowledge base based on graph similarity calculation. By statically analyzing the original code, a style feature vector is extracted based on code style indicator dimensions, wherein the code style indicator dimensions include code formatting rules, naming conventions, and design patterns; Based on the style feature vector, the best practice candidate solutions are style aligned to retain the structural characteristics of the original code and obtain the optimized solution.
[0111] In this embodiment, a historical version library management system is built to record each historical version of the project code and its associated refactoring information. This system is used to filter successful refactoring cases from the historical version library. The criteria for successful refactoring cases can be determined based on actual project requirements, such as improved performance, enhanced maintainability, and reduced defects in the refactored code.
[0112] Using static code analysis tools, we analyze the code before and after each successful refactoring case, generating a corresponding Abstract Syntax Tree (AST). An AST is a tree-like representation of code that clearly demonstrates the grammatical structure and logical relationships of the code. The differences between the ASTs before and after the code change are converted into graph samples. The nodes in the AST are treated as vertices, and the relationships between nodes (such as parent-child relationships and reference relationships) as edges. For the ASTs before and after the change, only the nodes and edges that differ are retained to construct the graph sample.
[0113] Design a graph neural network (GNN) model capable of processing graph-structured data. A GNN model typically consists of multiple graph convolutional layers, each of which aggregates and updates vertex information in the graph to learn a feature representation of the graph. The GNN model is trained using pre-constructed graph structure samples. During training, the graph structure samples are used as input, and the optimization effects of the reconstruction cases (such as performance improvement ratio and maintainability score) are used as labels. By continuously adjusting the model parameters, the model learns the mapping between graph structure and optimization effects. The target code is parsed to generate its abstract syntax tree (ABST). Simultaneously, the target code generated in the previous step is evaluated, which includes information on maintainability, security, and execution efficiency. The ABST and evaluation results of the target code are input into the trained GNN model. The model processes the ABST, extracts graph structure features, and combines the evaluation results to generate a comprehensive feature representation.
[0114] The historical knowledge base stores a large number of historical reconstruction cases and their corresponding graph structure samples. Using graph similarity calculation methods (such as those based on vertex and edge features), the similarity between the target code's graph structure features and each graph structure sample in the historical knowledge base is calculated. Based on the similarity calculation results, the top K best practice candidate solutions with the most similar graph structure to the target code are screened from the historical knowledge base. These candidate solutions share a certain degree of similarity with the target code in terms of code structure and optimization direction, potentially providing reference for optimizing the target code.
[0115] Clarify the dimensions of coding style indicators, including code formatting rules, naming conventions, and design patterns. Code formatting rules can include indentation, line breaking rules, and the use of brackets. Naming conventions can include naming conventions for variables, functions, and classes, such as using meaningful names and following camel case or underscore naming conventions. Design patterns can include the use of singleton, factory, and observer patterns in the code.
[0116] Use static code analysis tools to perform static analysis on the original code. Static analysis tools can scan code files and extract various information from the code, such as variable definitions, function calls, and class inheritance relationships. For example, by analyzing comments, spaces, line breaks, and other characters in the code, they can determine the formatting rules of the code; by analyzing the composition of variable and function names, they can identify naming conventions; and by analyzing the structure of classes and method call relationships, they can identify the use of design patterns.
[0117] Based on the determined code style indicator dimensions, a style feature vector is extracted from the static analysis results. A style feature vector is a multidimensional vector, with each dimension corresponding to a code style indicator, and the vector's value represents the characteristics of the original code with respect to that indicator. For example, for the code formatting rule dimension, multiple sub-dimensions can be set, such as indentation size and line break position, with each sub-dimension using a numerical value to represent the characteristics of the original code in that aspect. For the naming convention dimension, features such as the average length of variable and function names and whether specific naming prefixes are used can be counted.
[0118] The same static analysis is performed on the selected best practice candidates to extract the code style feature vector for each candidate. By comparing the style feature vectors of the candidate solutions with the style feature vectors of the original code, the differences between the candidate solutions and the original code in various code style indicators are analyzed.
[0119] Based on the differences in style feature vectors, the best practice candidates are style-aligned. The goal of style alignment is to preserve the structural features of the original code while aligning the candidate's code style with the original. For example, if the original code uses underscore naming and a candidate uses camelCase, the style alignment will convert the variable and function names in the candidate to underscore naming. If the original code uses four spaces for indentation and the candidate uses tabs, the candidate's indentation will be unified to four spaces.
[0120] After style alignment, an optimized solution is obtained that retains the design style of the original code. This optimization solution incorporates historical best practices and effectively optimizes the target code based on the evaluation results while maintaining readability and maintainability. For example, the optimization solution may refactor complex logic in the target code to adopt a more efficient design pattern, while maintaining the consistency of the code formatting, naming, and other styles with the original code, making it easier for team members to understand and maintain.
[0121] In this embodiment, based on the evaluation results, a graph neural network can be used to match historical best practices to generate an optimization solution that retains the original design style. This solution not only improves code quality and performance, but also reduces code maintenance costs, improves team development efficiency, and provides strong support for the sustainable development of software projects.
[0122] In some embodiments, in step S105, the static detection of the structural compliance of the abstract syntax tree of the optimization solution and the injection of test cases in the sandbox environment to verify the privacy leakage risk of the optimization solution during operation, and obtaining the verification results, specifically include: Parse the abstract syntax tree structure of the optimization solution, detect node type conflicts and structural violations based on the preset syntax rule set and style constraint library, and obtain structural compliance results; Generate specialized test case sets targeting privacy leakage risks through pre-built data flow graph analysis optimization solutions; Load the optimization solution code in an isolated environment and inject a specialized test case set; The code execution flow of the optimization solution in the isolated environment is captured through instrumentation technology, and the second sensitive information propagation path and resource consumption indicators are simultaneously recorded; The second sensitive information propagation path is matched and compared with the privacy rule base, and risk points that violate the data minimization principle or unauthorized output are marked to obtain the privacy leakage risk assessment results; Integrate structural compliance results and privacy leakage risk assessment results to generate a visual verification report that includes compliance levels and risk heat maps.
[0123] In this embodiment, a suitable static code analysis tool is selected that can parse code and generate an abstract syntax tree (AST). The optimization solution code is input into the tool, which performs lexical and syntactic analysis on the code, parsing the code layer by layer, and ultimately generating the corresponding abstract syntax tree. For example, for an optimization solution code containing function definitions, variable declarations, and loop statements, the abstract syntax tree will clearly display the function nodes, variable nodes, loop nodes, and the hierarchical relationships and connections between them.
[0124] Develop a detailed set of grammatical rules based on the grammatical specifications of the target programming language. This set of rules covers various grammatical structures of the language, such as the composition rules of statements, the legal form of expressions, and the rules of the type system. Establish a style constraint library based on the project's code style requirements. Style constraints include code indentation rules, naming conventions, comment formats, and more. Develop a structure detection module that can traverse and analyze the parsed abstract syntax tree. During the traversal process, check whether the type of each node complies with the rules and whether there are any violations in the structural relationships between nodes based on the preset grammatical rule set and style constraint library.
[0125] Use static code analysis tools and code annotations to construct a data flow diagram for the optimization solution. The data flow diagram records the flow of data from input to output within the code, including the data transfer relationships between variables, functions, and modules. Identify elements in the optimization solution that may contain private data, such as user personal information (name, ID number, contact information, etc.) and sensitive business data (transaction records, account balances, etc.). By analyzing variable names, function names, database table names, and other information in the code, combined with the data flow diagram, determine the source, processing, and output location of private data. Based on the analysis results of privacy-related elements, design a series of specialized test cases to address privacy leakage risks. These test cases should cover various possible privacy leakage scenarios, such as improper data storage, unauthorized access, and illegal data transmission. For example, design test cases to simulate hacker attacks to attempt to obtain private data stored in the optimization solution; or simulate illegal operations by internal personnel to verify whether they can bypass permission controls to access sensitive data. Combine the generated test cases into a specialized test case set.
[0126] Build a sandbox environment that is similar to the actual production environment. The environment should have independent hardware resources, operating system, runtime library, etc. The sandbox environment can be implemented through virtual machine technology or container technology to ensure that it is completely isolated from external systems to avoid affecting the actual system during the test process. Deploy the code of the optimization solution to the sandbox environment to ensure that the code can run normally in the sandbox environment. During the deployment process, necessary configuration and dependency installation are required to ensure that the code dependencies are complete and the versions are correct. Develop a test case injection tool that can automatically inject test cases from a specialized test case set into the optimization solution code running in the sandbox environment. Test cases can be provided to the code in a specific format (such as files, command line parameters, etc.) to trigger the code to execute the corresponding test logic.
[0127] Apply instrumentation technology to the optimization solution's code. Instrumentation involves inserting additional code at key locations to capture code execution information. Instrumentation can be performed at function call sites, variable assignment sites, data output sites, and other locations. When the optimization solution's code runs in a sandbox environment, the instrumentation code captures the code's execution flow in real time. This information includes the order of function calls, the number of loop executions, and the branch paths of conditional statements. The instrumentation code records this execution flow information in a log file for subsequent analysis. While capturing the code execution flow, the instrumentation code also records the secondary sensitive information propagation path. The secondary sensitive information propagation path refers to the actual flow of private data within the code during test case execution. By tracking the data transfer information recorded by the instrumentation code, the private data propagation path can be mapped. Resource consumption metrics, such as CPU usage, memory usage, and disk I / O, are collected during code execution. These metrics can be collected using system monitoring tools or custom monitoring code, and the collected data is recorded simultaneously with the execution flow information and the sensitive information propagation path.
[0128] Develop a privacy rule base that includes privacy-related rules such as the data minimization principle and authorized output rules. The data minimization principle requires that the code only collect and process necessary data, while the authorized output rules specify the conditions and methods for data output. Develop a risk assessment module that can read the recorded second-sensitive information propagation path and the privacy rule base. Match the second-sensitive information propagation path against the rules in the privacy rule base one by one to check whether the code violates the rules when processing private data. If the second-sensitive information propagation path is found to violate the rules in the privacy rule base, mark the corresponding risk point. The risk point should include information such as the specific location of the violation, the violated rule, and the severity of the violation. Summarize all marked risk points to form a privacy leakage risk assessment result.
[0129] Based on the structural compliance results, a set of compliance rating criteria is established. For example, compliance levels can be categorized into four levels: excellent, good, fair, and unsatisfactory, with ratings determined based on the number and severity of violations. If no violations are found during the structural compliance inspection, the rating is excellent; if a few minor violations are found, the rating is good, and so on.
[0130] Based on the risk point information from the privacy leakage risk assessment results, a risk heat map is drawn. This risk heat map uses different modules or functions in the code as dimensions, using different colors or brightness levels to indicate the level of risk. For example, high-risk modules or functions are indicated in red, while low-risk ones are indicated in green. This allows for a visual display of the distribution of privacy leakage risks in the code.
[0131] Integrate structural compliance results, compliance levels, risk heat maps, and privacy leakage risk assessment results to generate a visual verification report containing this information. This report can use a combination of charts, tables, and text descriptions to clearly display the verification results of the optimization solution, providing decision-making support for developers and testers.
[0132] In this embodiment, the optimization scheme's abstract syntax tree structure compliance can be accurately checked at a static level, effectively verifying the privacy leakage risks of the optimization scheme during runtime in a sandbox environment, and generating an intuitive visual verification report. This verification method can comprehensively and in-depth evaluate the quality and security of the optimization scheme, helping developers to promptly identify and resolve issues, thereby improving the quality and reliability of software development.
[0133] In some embodiments, in step S106, the step of comparing the abstract syntax tree differences between the original code and the verified optimization solution to replace only the functional logic nodes while retaining the legal code style features specifically includes: Create a dual-version abstract syntax tree of the original code and the optimized solution, and generate a node-level difference tree using a depth-first traversal algorithm; Based on the target location information of the code modification instruction set, the difference nodes are marked from the node-level difference tree; According to the node type and context relationship, the difference nodes are divided into functional nodes, style nodes and hybrid nodes; Replace the functional logic parts of the function nodes and the hybrid nodes, and retain the style features of the style nodes and the hybrid nodes; Parse the format rules and structural patterns of the original syntax tree in the dual-version abstract syntax tree to form a quantifiable style feature template; Apply style feature templates to rewrite the format of replaced function nodes and mixed nodes, perform syntax tree reconstruction to generate the final code file, and simultaneously generate change description documents.
[0134] In this embodiment, a depth-first traversal algorithm is used to traverse the abstract syntax trees of the original code and the optimization solution respectively. During the traversal process, each node is marked and recorded. For the corresponding nodes in the two trees, if there are differences in the type, attributes or sub-node structure of the nodes, these nodes are considered to be different. A node-level difference tree is generated, which uses the abstract syntax tree of the original code as the basic framework and records the node information in the optimization solution that is different from the original code in the difference tree. The nodes in the difference tree contain the location information, type information and specific difference content of the difference nodes. For example, the difference tree will record the differences between a function body in the optimization solution and the corresponding function body in the original code, including newly added statements, modified parameters, etc.
[0135] The code modification instruction set is generated during the code optimization process. It records the specific targets and location information of the optimization operations. For example, the instruction set may specify that a function needs to be optimized for performance, or that a variable needs to be renamed. Target positioning information is extracted from the code modification instruction set. This information includes the file path, line number, function name, variable name, etc. of the code, which is used to accurately locate the modification location in the code. The extracted target positioning information is matched with the node information in the node-level difference tree. For nodes that are successfully matched, they are marked in the difference tree, indicating that these nodes are difference nodes that need to be further processed. The markers can use specific identifiers or attributes so that they can be quickly identified in subsequent steps.
[0136] Perform type analysis on marked difference nodes, including function nodes, variable nodes, and expression nodes. Also, consider the node's context within the code, namely its connection and logical relationships with other nodes. For example, a variable node may be used within a function body. Its context includes the variable's definition location, usage location, and interactions with other variables within the function.
[0137] If a difference node primarily involves the functional logic of the code, such as algorithm implementation or business logic processing, and its modification has little impact on the code's style characteristics, then the node is classified as a functional node. For example, if a function that calculates the Fibonacci sequence is optimized and the calculation logic is modified, but the function's naming, comments, and other style characteristics remain unchanged, then the function node can be classified as a functional node.
[0138] If a difference node primarily involves code style features, such as code formatting (indentation, line breaks), naming conventions, and comment style, and has nothing to do with the code's functional logic, it is classified as a style node. For example, changing a variable name from camelCase to underscore in the code only affects the code style, and the corresponding node is a style node.
[0139] When a difference node involves both functional logic modifications and style feature changes, it is classified as a hybrid node. For example, if a function is optimized and its parameter names are modified, the function node is classified as a hybrid node.
[0140] For function nodes, simply replace the original code with the corresponding function node in the optimized solution. For hybrid nodes, extract the optimized solution's functional logic from the hybrid node and replace it with the corresponding functional logic in the original code. During the replacement process, ensure the correctness and integrity of the functional logic. For example, when replacing a function's functional logic, ensure that the function's input and output parameters, return value, and internal logical flow are consistent with the optimized solution.
[0141] For style nodes, their style features in the original code remain unchanged without any modification. For hybrid nodes, after replacing the functional logic, the style features of the original code in the hybrid node are retained. For example, the variable naming style and comment format in the hybrid node remain consistent with the original code.
[0142] Perform an in-depth analysis of the original code's abstract syntax tree to extract formatting rules and structural patterns. Formatting rules include code indentation (e.g., using four spaces for indentation), line breaks (e.g., wrapping after specific statements), and parentheses. Structural patterns include code module divisions and the organization of functions and classes. For example, analyze the definition formats of all functions in the original code to identify common structural patterns.
[0143] The extracted formatting rules and structural patterns are quantified and standardized to form a quantifiable style feature template. This template includes specific descriptions and quantitative indicators of various style features. For example, for indentation rules, the template stipulates that each level of indentation should use four spaces; for naming rules, the template stipulates that variable names should use a combination of lowercase letters and underscores, and class names should start with an uppercase letter.
[0144] Using the generated quantifiable style feature template, perform format checking and rewriting on the replaced function nodes and hybrid nodes. Check whether the nodes conform to the formatting rules and structural patterns in the template. If not, modify them to meet the template's requirements. For example, if the variable names in the replaced function nodes do not conform to the template's naming rules, rename them to conforming names.
[0145] After the format is rewritten, the modified abstract syntax tree is reconstructed. This process involves adjusting the node hierarchy and updating node reference information to ensure the structural correctness and integrity of the syntax tree. The reconstructed abstract syntax tree is then converted back into code to generate the final code file.
[0146] During the code generation process, all modification operations are recorded, including node replacements, retention and rewriting of stylistic features, and other information. Based on these records, a change log is generated, detailing the differences between the original and optimized code, including functional logic modifications and retention of stylistic features. This change log helps developers and maintainers understand the code changes and facilitates subsequent code maintenance and upgrades.
[0147] This example accurately compares the abstract syntax tree differences between the original code and the optimized solution, replacing only functional logic nodes while preserving valid coding style features. This method optimizes code while ensuring consistent coding style, improving code readability and maintainability, and reducing code maintenance costs and the difficulty for team members to understand.
[0148] Reference Figure 2 An embodiment of the present invention provides a code reconstruction system 2 based on function identification, and the system 2 specifically includes: The structure analysis module 201 is used to scan the project file topology structure and build a knowledge graph including API call chains and data flow paths based on code import relationships and technology stack features; Function identification module 202, used to parse the code function description input by the user through a pre-trained natural language processing model, and perform context enhancement in combination with the knowledge graph to generate a code modification instruction set with weighted annotations; The quantitative evaluation module 203 is used to perform a three-dimensional quantitative evaluation of the maintainability, security, and execution efficiency of the target code based on the code modification instruction set, combined with data flow sensitivity marking and performance profile analysis, to obtain an evaluation result; The code optimization module 204 is used to match historical best practice patterns through a graph neural network based on the evaluation results and generate an optimization solution that retains the original design style; The code verification module 205 is used to detect the structural compliance of the abstract syntax tree of the optimization solution at the static level, and inject test cases into the sandbox environment to verify the privacy leakage risk of the optimization solution during operation, and obtain verification results; The code replacement module 206 is configured to replace only functional logic nodes and retain legal code style features by comparing the abstract syntax tree differences between the original code and the verified optimization solution.
[0149] It is understandable that if Figure 1 The contents of the embodiment of the code reconstruction method based on function identification shown in the figure are applicable to the embodiment of the code reconstruction system based on function identification. The functions specifically implemented by the embodiment of the code reconstruction system based on function identification are similar to those in the embodiment of the code reconstruction method based on function identification shown in the figure. Figure 1 The embodiment of the code reconstruction method based on function identification shown in FIG. 1 is the same as that in FIG. 1 and the beneficial effects achieved are the same as those in FIG. Figure 1 The beneficial effects achieved by the embodiment of the code reconstruction method based on function identification are also the same.
[0150] It should be noted that the information interaction, execution process and other contents between the above-mentioned systems are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0151] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0152] Reference Figure 3 An embodiment of the present invention further provides a computer device 3, comprising: a memory 302, a processor 301, and a computer program 303 stored in the memory 302. When the computer program 303 is executed on the processor 301, the code reconstruction method based on function identification as described in any one of the above methods is implemented.
[0153] The computer device 3 may be a desktop computer, a notebook computer, a PDA, a cloud server or other computing devices. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 This is merely an example of the computer device 3 and does not constitute a limitation on the computer device 3 . The computer device 3 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device 3 may also include input and output devices, network access devices, etc.
[0154] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0155] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard drive or memory of the computer device 3. In other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device 3. Furthermore, the memory 302 may include both an internal storage unit of the computer device 3 and an external storage device. The memory 302 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 302 may also be used to temporarily store data that has been output or is about to be output.
[0156] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the code reconstruction method based on function identification as described in any one of the above methods is implemented.
[0157] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0158] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0159] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0160] In the embodiments disclosed in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0161] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A code refactoring method based on function identification, characterized in that: The method specifically includes: Scan the project file topology and build a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features; The pre-trained natural language processing model parses the user-entered code function description, combines it with the knowledge graph for context enhancement, and generates a weighted annotated code modification instruction set. Based on the code modification instruction set, combined with data flow sensitivity marking and performance profiling analysis, a three-dimensional quantitative evaluation of the target code in terms of maintainability, security, and execution efficiency is performed to obtain the evaluation results; Based on the evaluation results, a graph neural network is used to match historical best practice patterns and generate an optimization solution that retains the original design style. Static testing of the structural compliance of the optimization scheme's abstract syntax tree is performed. Test cases are injected into the sandbox environment to verify the privacy leakage risks of the optimization scheme during runtime and obtain verification results. By comparing the abstract syntax tree differences between the original code and the verified optimized solution, only functional logic nodes are replaced and legal code style features are retained.
2. The method according to claim 1, characterized in that The scanning of the project file topology structure builds a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features, specifically including: Analyze the project file topology and establish a module hierarchy index based on file path dependencies; Identify core entities from the source code based on the module level index, wherein the core entities include API definitions, data objects, and external dependencies; Based on the code call stack and data flow analysis, a four-dimensional relationship network including call, inheritance, transmission, and dependency is established between core entities; According to the technology stack declaration in the project configuration file, framework constraint rules and best practice templates are injected into the four-dimensional relationship network to obtain an enhanced four-dimensional relationship network; The enhanced four-dimensional relationship network is converted into an attribute graph structure, wherein the nodes of the attribute graph structure are used to carry code location identifiers, and the edges of the attribute graph structure are used to mark relationship types and call frequencies.
3. The method according to claim 1, characterized in that The pre-trained natural language processing model is used to parse the user's input code function description, and combined with the knowledge graph for context enhancement, to generate a code modification instruction set with weighted annotations, specifically including: A Transformer-based bidirectional encoder is used to parse the functional description input by the user, and a code function semantic vector is generated through the semantic encoding layer; Based on the code function semantic vector, match related entities and relationship subgraphs from the knowledge graph; Using the retrieved subgraph structure, we perform graph attention weighting on the code function semantic vector to generate a context-enhanced semantic vector. Parsing an atomic operation sequence according to the context-enhanced semantic vector, where each atomic operation in the atomic operation sequence includes a target code location and a modification type; Based on the entity sensitivity and call frequency in the knowledge graph, the execution priority weights are assigned to the atomic operations to obtain a weighted atomic operation sequence; Encapsulate weighted atomic operation sequences into a machine-executable code modification instruction set.
4. The method according to claim 1, wherein Based on the code modification instruction set, combined with data flow sensitivity marking and performance profiling analysis, the target code is subjected to a three-dimensional quantitative evaluation of maintainability, security, and execution efficiency, and the evaluation results are obtained, including: Extract the code segment to be evaluated based on the target code location and operation type in the code modification instruction set; Track the first sensitive information propagation path of the code segment to be evaluated through a pre-built data flow graph, and calculate the security risk score based on the first sensitive information propagation path using the privacy rule base; Execute the code segment to be evaluated in the sandbox environment, collect resource consumption indicators, establish an execution efficiency profile model, and output the execution efficiency profile; Based on the code complexity index and change history data, the maintainability prediction model is used to analyze the code segment to be evaluated and generate the evolution cost coefficient; The code complexity indicators include cyclomatic complexity, change cost, and dependency entropy. The cyclomatic complexity is calculated by the McCabe index, the change cost is calculated by multiplying the file modification frequency and the impact range, and the dependency entropy is calculated by multiplying the number of external dependencies and the version dispersion. The security risk score, execution efficiency profile, and evolution cost coefficient are input into the evaluation matrix to generate a three-dimensional evaluation result including a weight vector.
5. The method according to claim 1, wherein Based on the evaluation results, the graph neural network is used to match historical best practice patterns to generate an optimization solution that retains the original design style. Specifically, it includes: Extract successful refactoring cases based on the historical version library and convert the abstract syntax tree differences before and after the code changes of the successful refactoring cases into graph structure samples; A graph neural network model is built based on the graph structure sample. The abstract syntax tree of the target code and the evaluation results are input into the graph neural network model for matching. The top K best practice candidate solutions are screened from the historical knowledge base based on graph similarity calculation. By statically analyzing the original code, a style feature vector is extracted based on code style indicator dimensions, wherein the code style indicator dimensions include code formatting rules, naming conventions, and design patterns; Based on the style feature vector, the best practice candidate solutions are style aligned to retain the structural characteristics of the original code and obtain the optimized solution.
6. The method according to claim 1, characterized in that The structural compliance of the abstract syntax tree of the optimization solution is detected at the static level, and test cases are injected into the sandbox environment to verify the privacy leakage risk of the optimization solution during operation, and the verification results are obtained, including: Parse the abstract syntax tree structure of the optimization solution, detect node type conflicts and structural violations based on the preset syntax rule set and style constraint library, and obtain structural compliance results; Generate specialized test case sets targeting privacy leakage risks through pre-built data flow graph analysis optimization solutions; Load the optimization solution code in an isolated environment and inject a specialized test case set; The code execution flow of the optimization solution in the isolated environment is captured through instrumentation technology, and the second sensitive information propagation path and resource consumption indicators are simultaneously recorded; The second sensitive information propagation path is matched and compared with the privacy rule base, and risk points that violate the data minimization principle or unauthorized output are marked to obtain the privacy leakage risk assessment results; Integrate structural compliance results and privacy leakage risk assessment results to generate a visual verification report that includes compliance levels and risk heat maps.
7. The method according to any one of claims 1 to 6, characterized in that By comparing the abstract syntax tree differences between the original code and the verified optimization solution, only functional logic nodes are replaced while retaining legal code style features, specifically including: Create a dual-version abstract syntax tree of the original code and the optimized solution, and generate a node-level difference tree using a depth-first traversal algorithm; Based on the target location information of the code modification instruction set, the difference nodes are marked from the node-level difference tree; According to the node type and context relationship, the difference nodes are divided into functional nodes, style nodes and hybrid nodes; Replace the functional logic parts of the function nodes and the hybrid nodes, and retain the style features of the style nodes and the hybrid nodes; Parse the format rules and structural patterns of the original syntax tree in the dual-version abstract syntax tree to form a quantifiable style feature template; Apply style feature templates to rewrite the format of replaced function nodes and mixed nodes, perform syntax tree reconstruction to generate the final code file, and simultaneously generate change description documents.
8. A code reconstruction system based on function identification, characterized in that: The system specifically includes: The structure analysis module is used to scan the project file topology and build a knowledge graph containing API call chains and data flow paths based on code import relationships and technology stack features; Function recognition module, which uses a pre-trained natural language processing model to parse the user-entered code function description, combines it with the knowledge graph for context enhancement, and generates a weighted annotated code modification instruction set; The quantitative evaluation module is used to perform a three-dimensional quantitative evaluation of the target code in terms of maintainability, security, and execution efficiency based on code modification instruction sets, combined with data flow sensitivity tagging and performance profiling analysis, to obtain evaluation results. The code optimization module is used to match historical best practice patterns based on evaluation results through graph neural networks to generate optimization solutions that retain the original design style; The code verification module is used to check the structural compliance of the abstract syntax tree of the optimization solution at the static level, and inject test cases into the sandbox environment to verify the privacy leakage risk of the optimization solution during operation and obtain verification results; The code replacement module is used to replace only functional logic nodes and retain legal code style features by comparing the abstract syntax tree differences between the original code and the verified optimized solution.
9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory, which implements the code reconstruction method based on function identification according to any one of claims 1 to 7 when the computer program is executed on the processor.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the code reconstruction method based on function identification as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Distributed environment-oriented code security detection and version control method and system
CN121277544A
Code security detection and version control method and system for distributed environment
CN121277544B
API knowledge graph construction method, medium and equipment
CN121501339A
Automatic file processing-oriented intelligent retrieval enhancement instruction generation system
CN122220308A
Historical engineering code parsing optimization method and device, equipment and storage medium
CN122387506A