Automated Code Activity Diagram Generation System and Method
Through the automated code activity diagram generation system, the source code analysis module, model conversion engine, model warehouse, activity diagram generator and layout module are used to solve the problem of manual drawing of activity diagrams that consume time and effort and cannot take into account the needs of multiple angles, and efficient and multi-angle activity diagram generation is achieved, improving development efficiency and system visibility.
Patent Information
- Application Number
- CN202510193754.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-21
AI Technical Summary
In the prior art, manual drawing of activity diagrams is time-consuming and labor-intensive, difficult to update in a timely manner, and cannot take into account the needs of technical personnel and business personnel, and cannot provide multi-angle process display.
Develop an automated code activity diagram generation system, including source code parsing module, model conversion engine, model warehouse, activity diagram generator and layout module, and automatically parsing source code, converting it into an abstract syntax tree, generating programming language abstract models and generating activity diagrams based on different perspectives.
It realizes automatic generation of activity diagrams that match the actual code logic, reduces the time and manpower of manual drawing, improves development efficiency, and can generate activity diagrams required by technical personnel and business personnel at the same time, improving the visibility and understanding of the system.
Smart Images

Figure CN119690440B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to an automated code activity diagram generation system and method. Background Art
[0002] In the prior art, as an important tool in system development, activity diagrams are usually used to display the internal operation processes of a system. To help designers and developers clarify the logical relationships of the system, specialized drawing tools are often used to manually draw activity diagrams during the development process. This method is widely applied in various software development projects. Especially in the system design phase, activity diagrams help to display the logical structure of the system and the method call relationships, and help the development team and project managers intuitively understand the business and technical processes of the system.
[0003] However, the method of manually drawing activity diagrams in the prior art has some obvious deficiencies. First, activity diagrams are usually drawn before development, and after the system development is completed, the activity diagrams are often not updated in a timely manner, resulting in differences between the design diagrams and the actual implemented code, and unable to accurately reflect the actual running logic of the current system. Second, manually drawing activity diagrams requires a lot of time and effort. Especially for systems with complex functions or frequent iterations, manually maintaining activity diagrams becomes very cumbersome. In addition, the activity diagrams generated by existing tools are usually only for technical personnel, and it is difficult to take into account the needs of business personnel, and unable to provide multi-angle process displays.
[0004] In view of the above problems, there is an urgent need to develop a new code activity diagram generation system. Summary of the Invention
[0005] The present application provides an automated code activity diagram generation system and method to improve the efficiency of developers.
[0006] The automated code activity diagram generation system provided by the present application includes:
[0007] A source code parsing module, configured to extract source code from a code file and generate an abstract syntax tree corresponding to the source code;
[0008] A model conversion engine, operating based on query view conversion specifications and configured with a set of parsing rules, for parsing the method call relationships and statement execution orders in the abstract syntax tree, and converting the source code into a programming language abstract model;
[0009] A model repository, configured to store the programming language abstract model and identify the programming language type according to the extension name of the source code file to match the corresponding programming language abstract model;
[0010] An activity diagram generator, connected to the model repository, is used to generate activity diagrams according to the programming language abstraction model. The activity diagrams include activity diagrams generated based on logical code and activity diagrams generated based on code comments.
[0011] A layout module is used to determine the positions, sizes, and shapes of each node in the activity diagram and generate connection lines between the nodes by parsing the method call relationships and code execution paths in the programming language abstraction model.
[0012] Furthermore, the source code parsing module jointly uses a lexical analyzer and a syntax analyzer to perform lexical analysis on the source code, extract keywords, identifiers, and symbols in the source code; and construct an abstract syntax tree through syntax analysis.
[0013] Furthermore, the source code parsing module is used to parse the source code of multiple programming languages, including Java, Python, and C++, and dynamically adjust the parsing process according to the syntax rules of the programming language.
[0014] Furthermore, the model repository stores the programming language abstraction model through a hash table; and uses the extension name and unique identifier of the source code file as the hash key to quickly retrieve the abstraction models of different programming languages.
[0015] Furthermore, the model repository includes a version management unit, which is used to store and retrieve the programming language abstraction models of multiple versions of the same source code file and compare the differences between different versions.
[0016] Furthermore, the activity diagram generator includes a multi-perspective activity diagram generation unit, which is used to generate activity diagrams from multiple perspectives according to different user requirements. Among them, the perspectives include:
[0017] The perspective of developers, which is used to display the detailed code execution logic;
[0018] The perspective of business personnel, which is used to generate a simplified business logic activity diagram according to code comments;
[0019] The perspective of historical versions, which is used to display the function iteration or code modification history of the system.
[0020] Furthermore, the activity diagram generator converts each logical node in the programming language abstraction model into a visual activity diagram node through a graphics rendering engine and draws the connection lines between the nodes according to the function call relationship.
[0021] Further, the layout module includes a graph optimization unit, which adjusts the node layout of the activity graph based on a multi-objective optimization algorithm to reduce the distance deviation between nodes, reduce the number of crossings of connection lines, and optimize the alignment degree of nodes. The multi-objective optimization algorithm is solved by a constraint optimization model provided by the following formula (1):
[0022]
[0023] where represents the logical connection weight between the -th and the -th nodes, and the logical connection weight is determined based on the call relationship and dependency degree of the nodes; represents the Euclidean distance between the -th and the -th nodes; is the ideal distance determined according to the logical relationship; represents the length of the -th connection line; is the length of the longest connection line in the graph; represents the number of crossings of the -th connection line; represents the alignment deviation of the -th node; is the number of connection lines with the most crossings in the graph; is the maximum alignment deviation in the current graph; represents the distance between the -th node and the boundary of the activity graph; is the minimum boundary distance among all nodes; represents the number of nodes; represents the number of connection lines; , , and are trade-off parameters; is an adjustment parameter for balancing the influence of the connection line length and the number of crossings.
[0024] Further, the automated code activity graph generation system further includes a code processing flow complexity adjustment unit, which is used to calculate the complexity of the code processing flow according to the abstract syntax tree generated from the source code, and adjust the number of nodes of the activity graph according to the code processing complexity. Among them, the calculation formula of the code processing complexity is calculated by the following formula (2):
[0025]
[0026] where Indicates the complexity of the code processing flow; Indicates the number of execution statements in the th code snippet; Indicates the branch depth in the th code snippet; Indicates the call chain length of the th function; Indicates the dependency depth of the th function; Indicates the number of parameters of the and respectively indicate the number of code snippets and functions;
[0027] The code processing flow complexity adjustment unit adjusts the number of nodes and the level of detail of the activity diagram based on the following rules:
[0028] When the calculated code processing complexity exceeds a preset first threshold, the details inside the function in the activity diagram are merged into a single node, and only the high-level logic of the function call order is retained;
[0029] When the calculated code processing complexity exceeds a preset second threshold, multiple related code snippets in the activity diagram are merged into a single node; wherein, the second threshold is greater than the first threshold;
[0030] When the calculated code processing complexity is lower than a preset third threshold, the detailed execution logic inside the function in the activity diagram is displayed, including branches, loops, and specific operation steps.
[0031] This application provides an automated code activity diagram generation method, including:
[0032] Extract the source code from the code file and generate an abstract syntax tree corresponding to the source code;
[0033] Based on the query view transformation specification operation, using a set of configured parsing rules, parse the method call relationships and statement execution orders in the abstract syntax tree, and convert the source code into a programming language abstract model;
[0034] Store the programming language abstract model in the model repository, and identify the programming language type according to the extension of the source code file to match the corresponding programming language abstract model;
[0035] Through the model repository, generate an activity diagram based on the programming language abstract model, where the activity diagram includes an activity diagram generated based on logical code and an activity diagram generated based on code comments;
[0036] By analyzing the method call relationships and code execution paths in the programming language abstraction model, the positions, sizes, and shapes of the nodes in the activity diagram are determined, and the connection lines between the nodes are generated.
[0037] The beneficial effects of this application mainly include: (1) Through source code parsing and model conversion, the system can automatically generate an activity diagram that matches the actual code logic, avoiding the deviations that may occur when manually drawing the activity diagram. This system can reflect the actual processing flow of the code in real time, helping developers and relevant personnel understand the operation logic of the system more accurately and perform effective process optimization. (2) By automatically generating the activity diagram, this system reduces the time and manpower required for manually drawing the activity diagram. Developers can focus on writing code, and the activity diagram is automatically generated after the system development is completed, quickly and accurately showing the code flow, greatly improving the development efficiency, especially suitable for systems with complex functions or long-term iterations. (3) The system can generate two different types of activity diagrams simultaneously, meeting the needs of technical personnel and business personnel respectively. The activity diagram generated based on the logical code helps developers understand the code flow, while the activity diagram generated based on the code comments helps business personnel master the business logic of the system, enhancing the visibility and understandability of the system. Description of the Drawings
[0038] Figure 1 is a schematic diagram of an automated code activity diagram generation system provided by the first embodiment of this application.
[0039] Figure 2 is a flowchart of an automated code activity diagram generation method provided by the second embodiment of this application. Detailed Embodiments
[0040] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific embodiments disclosed below.
[0041] The first embodiment of this application provides an automated code activity diagram generation system. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of this application. The following will be described in detail an automated code activity diagram generation system provided by the first embodiment of this application in conjunction with Figure 1 .
[0042] The automated code activity diagram generation system includes a source code parsing module 101, a model conversion engine 102, a model repository 103, an activity diagram generator 104, and a layout module 105.
[0043] The source code parsing module 101 is used to extract the source code from the code file and generate an abstract syntax tree corresponding to the source code.
[0044] The source code parsing module 101 plays a crucial role in the automated code activity diagram generation system. The main function of this module is to extract the code content from the source code file and, through a series of analysis processes, generate an abstract syntax tree (AST, Abstract Syntax Tree) corresponding to the source code structure. The workflow of this module generally includes the following steps.
[0045] First of all, the source code parsing module 101 can support the parsing of multiple programming languages, including but not limited to Java, Python, C++, etc. To achieve the support for multiple languages, the source code parsing module 101 includes a language recognition unit, which automatically determines the programming language type of the code to be parsed by analyzing the file extension or file header information of the source code file. This multi-language adaptability provides great flexibility for the system and can be widely applied to code parsing and activity diagram generation in different language environments.
[0046] Next, the source code parsing module 101 starts to perform lexical analysis and syntactic analysis on the source code. In the lexical analysis stage, the parsing module 101 disassembles the source code into basic language units, called lexical tokens (Tokens), which include keywords, operators, identifiers, constants, etc. Through lexical analysis, the basic structure of the source code is extracted and further processed.
[0047] Subsequently, the source code parsing module 101 parses the sequence of lexical tokens through a parser to construct an abstract syntax tree (AST). The AST is a tree-like data structure that can reflect the hierarchical structure and syntactic relationships of the source code. Each node represents a structure in the source code, such as variable declarations, function definitions, conditional judgments, or loops. The parser adopts a rule set based on context-free grammar to ensure the accuracy of parsing and the integrity of the code structure. Specifically, the parser checks the syntactic correctness of the code and parses the code step by step according to the preset syntactic rules. If an error is detected during the syntactic analysis process, the source code parsing module 101 also has an error reporting function, which can output detailed error information to help developers quickly locate and correct problems in the code.
[0048] After generating the abstract syntax tree, the source code parsing module 101 further optimizes the structure of the AST. This process involves simplifying and standardizing the AST, such as removing irrelevant syntax elements (such as blank lines or invalid comments), merging certain redundant structures, while keeping the logical structure of the code unchanged. This optimization step not only helps improve the efficiency of subsequent model transformation, but also ensures that the process of generating the activity diagram is more efficient and accurate.
[0049] To improve the speed of code parsing, the source code parsing module 101 adopts a parallel parsing mechanism. For large code files or multi-file projects, the parsing module can process multiple source code files simultaneously, allocating multiple threads or processes to perform lexical analysis and syntax analysis. This parallel processing mechanism greatly improves the parsing efficiency, especially suitable for the parsing process of large projects or complex code libraries.
[0050] In addition to parsing the structure of the code, the source code parsing module 101 can also extract relevant comment information from the source code. Comment information often contains descriptions related to business logic or additional developer annotations, and this information can provide auxiliary basis when generating the activity diagram later. The source code parsing module 101 can identify different types of comments in the code (such as single-line comments, multi-line comments, and documentation comments), and combine this information with the abstract syntax tree of the code to ensure that the mapping relationship between comments and code logic is reflected when generating the activity diagram.
[0051] In practical applications, the source code parsing module 101 works closely with other modules of the system. For example, the abstract syntax tree generated by the parsing module is directly passed to the model transformation engine 102, which further converts the AST into an abstract model of the programming language according to predefined rules. This process highly depends on the structured data provided by the source code parsing module 101 to ensure that the logical relationship of the code is accurately reflected in subsequent steps.
[0052] Generally speaking, the source code parsing module 101 is not only one of the basic modules of this system, but also ensures the accurate extraction and optimization of the source code structure through its support for multiple languages, parallel parsing mechanism, comment extraction, and error reporting functions. This module provides a stable and efficient foundation for subsequent model transformation and activity diagram generation, and can adapt to different development environments and code scales. It is the core component for realizing the automatic generation of code activity diagrams.
[0053] Furthermore, the source code parsing module, through the combined action of the lexical analyzer and the syntax analyzer, performs lexical analysis on the source code, extracts keywords, identifiers, and symbols in the source code; and constructs an abstract syntax tree through syntax analysis.
[0054] First, the main task of the lexical analyzer is to break down the character sequence in the source code into lexical units, which are the basic elements in a programming language, such as keywords, identifiers, operators, constants, and delimiters. The lexical analyzer scans the source code file, identifies these basic lexical elements, and assigns corresponding categories to each element. For example, 'if' is recognized as a keyword for conditional judgment, 'x' is recognized as an identifier, and '+' is recognized as an operator. The result of lexical analysis is a list composed of lexical units, which retains the logical order of the source code and provides structured data input for the subsequent syntax analyzer.
[0055] The process of lexical analysis is not just a simple character splitting of the source code. The lexical analyzer must be able to recognize the syntax rules of the programming language to ensure that identifiers and keywords can be accurately distinguished. In addition, the lexical analyzer also needs to handle comments and whitespace characters in the source code. Usually, comments do not affect the execution of the code, so the lexical analyzer will recognize and mark comments as ignorable elements, while handling whitespace symbols (such as spaces, newlines, etc.) in the code to ensure the integrity of the code structure. During this process, the lexical analyzer also has a certain degree of fault tolerance. If illegal characters or format errors are found in the source code during parsing, the system will generate error messages in a timely manner to facilitate developers to make corrections.
[0056] Next, the list of lexical units generated by the lexical analyzer will be passed to the syntax analyzer. The syntax analyzer parses the lexical units structurally according to the syntax rules of the programming language to construct the corresponding abstract syntax tree (AST) of the source code. The abstract syntax tree is a tree-like structure, and each node represents a syntactic structure unit in the code, such as an expression, a statement, a function definition, or a control structure. The syntax analyzer gradually constructs this tree by analyzing the order and nesting relationship of the lexical units, so that it can reflect the hierarchical structure of the source code. For example, the conditional judgment and execution body of an 'if' statement will be parsed as sub-nodes of the 'if' node, while the control condition and loop body of a loop statement will form an independent loop structure node.
[0057] The role of the syntax analyzer is not only to transform the source code into an abstract syntax tree, but also to ensure the syntactic correctness of the code. By defining and following the context-free syntax rules of the programming language, the syntax analyzer can check whether there are syntax errors in the code. For example, it can identify problems such as unmatched parentheses, incomplete statement structures, or undefined function calls. If the syntax analyzer finds an error in the code, it will generate corresponding error messages and indicate the location and type of the error to help developers quickly locate and fix the problem.
[0058] During the construction of the abstract syntax tree, the parser can also perform certain optimizations and simplifications on the code structure. For example, for some redundant expressions or unnecessary nested structures, the parser can optimize them during tree construction, thereby simplifying the code structure and providing a more efficient data basis for subsequent model transformation. This optimization process does not change the logic of the source code but improves the parsing efficiency of the system by streamlining the code structure.
[0059] Finally, the combined action of the lexical analyzer and the parser generates a complete abstract syntax tree. This tree not only preserves the logical structure of the source code but also presents elements such as keywords, identifiers, and operators in the code in a structured form, providing the necessary basic data for the next steps of the system, such as the generation of the programming language abstraction model and the creation of activity diagrams. Those skilled in the art can easily implement the specific functions of the source code parsing module according to the above detailed description through existing lexical analysis and parsing techniques.
[0060] Furthermore, the source code parsing module is used to parse the source code of multiple programming languages, including Java, Python, and C++, and dynamically adjust the parsing process according to the syntax rules of the programming language.
[0061] When the source code parsing module receives a source code file, the system first automatically identifies the programming language type of the file through the file extension or the characteristic information at the file header. For example, the extension of a Java file is usually ".java", a Python file is ".py", and a C++ file may be ".cpp" or ".h". Through this file type recognition mechanism, the system can quickly determine which type of parsing rules need to be applied. Next, the source code parsing module will dynamically load the lexical analysis and parsing rule sets corresponding to the identified programming language type.
[0062] For each programming language, the syntax and structure of the source code are different. Take Java as an example. Java is an object-oriented programming language, and its source code structure includes classes, interfaces, methods, fields, and access modifiers, etc. Therefore, when parsing Java code, the source code parsing module needs to identify and process these specific language elements to construct an abstract syntax tree that conforms to the characteristics of the Java language. On the contrary, for an interpreted language like Python, due to its more flexible and dynamic syntax, when the source code parsing module processes Python code, it needs to pay attention to characteristics such as the indentation hierarchy structure and the definition of dynamically typed variables.
[0063] As a more complex compiled language, the C++ language not only includes object-oriented features but also supports advanced language features such as multiple inheritance and template metaprogramming. Therefore, when the source code parsing module processes C++ code, it must not only parse basic class, function, and variable definitions but also handle templates, macro definitions, and syntax rules related to memory management. To this end, the system adjusts the parsing depth according to the complexity of C++, ensuring that all language features can be correctly identified and parsed.
[0064] Another feature of the source code parsing module is its ability to dynamically adjust the parsing process. The code structures of different programming languages vary, and when dealing with these differences, the parsing module can automatically adjust the parsing strategy according to the specific rules of the language. For example, function and method calls in Java and C++ are based on static types, so the lexical analyzer and syntax analyzer can parse based on a clear type system. Since Python is a dynamically typed language and its variable types are determined at runtime, when parsing Python code, the source code parsing module needs to adopt a more flexible strategy to handle the dynamic binding of variables and function calls. By dynamically adjusting the parsing strategy, the system can accurately parse various types of programming languages without sacrificing efficiency.
[0065] In addition, the source code parsing module also supports the parsing of cross-language code. In modern software development, projects often consist of a mixture of multiple programming languages. For example, Java and Python are used simultaneously in a system to handle different business logics. To adapt to this multi-language development environment, the source code parsing module has the ability to co-parse multiple languages. When processing different code files of the same project, it can automatically switch the language rule set and uniformly manage the parsing results of each language. This multi-language support provides great convenience for developers and ensures the compatibility and extensibility of the system in complex environments.
[0066] Through this dynamic adjustment of the parsing process based on programming language features, the source code parsing module can accurately process multiple programming languages and generate an abstract syntax tree that conforms to the characteristics of each language, laying the foundation for the subsequent generation of activity diagrams. The design of this module is not only applicable to common programming languages but also extensible, making it easy to add support for other languages.
[0067] The model transformation engine 102 operates based on the query view transformation specification and is configured with a set of parsing rules for parsing the method call relationships and statement execution order in the abstract syntax tree, converting the source code into a programming language abstract model.
[0068] The model transformation engine 102 is a key component in the automated code activity diagram generation system of the present invention. Its function is to convert the abstract syntax tree (AST) generated from the source code parsing module 101 into an abstract model of the programming language for further generation of the activity diagram. This engine is based on the Query / View / Transformation (QVT), which is a standardized framework for model transformation and can achieve the transformation between different models, especially playing an important role in the process of transforming from the abstract syntax tree to the abstract model.
[0069] The core of the model transformation engine 102 lies in a set of parsing rules it configures. These rules define how to extract the structural information related to method call relationships and statement execution orders from the abstract syntax tree. Each parsing rule operates on specific programming language structures, such as class definitions, function calls, control flow statements (such as conditional judgments, loops), etc. Through these rules, the model transformation engine 102 can identify the key elements in the source code and map them to the abstract model of the programming language. This engine not only supports basic language structures but also can handle advanced language features such as complex nested calls and recursive calls, thus ensuring the integrity and accuracy of the code structure.
[0070] In the specific operation process, the model transformation engine 102 first receives the abstract syntax tree generated by the source code parsing module. The nodes in the abstract syntax tree represent the various language structures in the source code, and these nodes are organized according to the grammar rules to form the hierarchical structure of the code. The model transformation engine first identifies the root node in the abstract syntax tree and processes each child node in turn. For each node, the model transformation engine extracts the information in the node according to the corresponding parsing rule. For example, if a node represents a function call, the transformation engine will extract the name of the function, its parameters, and the return value type of the call, and generate the corresponding function call object in the abstract model of the programming language. In this way, the abstract syntax tree is gradually converted into the abstract model of the programming language.
[0071] The model transformation engine 102 can handle multiple programming languages. Therefore, its parsing rules can be adjusted according to the grammar characteristics of different languages. For example, for object-oriented languages (such as Java or C++), the engine pays attention to class inheritance relationships, interface implementations, and polymorphism handling; while for scripting languages (such as Python), it focuses more on functional programming features and dynamic type resolution. Therefore, the model transformation engine is configured with a language adaptation module that can identify the language type of the source code and select the corresponding set of parsing rules according to the specific grammar and structure of the language, thus ensuring that each programming language can be accurately processed.
[0072] In addition to syntax parsing, the model transformation engine 102 also has a semantic analysis function. Semantic analysis is a deeper understanding of code logic that goes beyond surface-level syntax rules. For example, when dealing with variable scopes, the transformation engine needs to determine the lifespan of variables and their scope within different functions and classes to ensure logical consistency during activity diagram generation. Through semantic analysis, the model transformation engine can understand the dependencies in the code, the call chains between methods, and the interaction between global and local variables. This process is crucial for ensuring that the generated activity diagram accurately reflects the code logic.
[0073] Furthermore, the model transformation engine 102 fully considers the dynamic characteristics of the code. For example, in some cases, the source code may contain features such as conditional compilation, dynamic type binding, or runtime reflection, which are difficult to statically analyze at compile time. To address this issue, the model transformation engine is equipped with a dynamic analysis module that can, when necessary, simulate the code's runtime environment to obtain the actual execution paths of the code under different conditions. In this way, the transformation engine can generate a more complete abstract model, thereby ensuring the accuracy of the activity diagram.
[0074] The model transformation engine also supports incremental transformation. During actual development, the source code may be frequently updated or modified, and re-parsing the entire code from scratch and generating an activity diagram may consume a large amount of time. To this end, the model transformation engine can identify the changed parts of the source code and only re-parse and transform the modified code snippets without having to perform a full parse of the entire codebase. This incremental transformation mechanism greatly improves the efficiency of the system, especially for large codebases and complex systems.
[0075] After the abstract model is generated, the model transformation engine 102 stores the generated model in the model repository 103 for further use by the activity diagram generator 104. To ensure the efficiency of data transmission and storage, the model transformation engine also has model compression and optimization functions that can reduce the size of the abstract model by eliminating redundant nodes and simplifying data structures, thereby accelerating the activity diagram generation process.
[0076] Generally speaking, the model transformation engine 102 is one of the core components of this system for realizing automated code activity diagram generation. Through flexible parsing rules, multi-language support, semantic analysis, and incremental transformation, it ensures that the complex structure of the source code can be accurately transformed into an abstract model of the programming language, providing a solid foundation for subsequent activity diagram generation.
[0077] The model repository 103 is used to store the programming language abstract model and identify the programming language type based on the extension of the source code file to match the corresponding programming language abstract model.
[0078] The model repository 103 is an important component responsible for storing and managing the programming language abstract models generated by the model conversion engine 102. Its role is not only a simple storage function, but also includes classifying, indexing, retrieving, and version management of the stored models. The model repository plays a crucial role in connecting the upper and lower levels of the entire system because the activity diagram generator 104 depends on extracting abstract model data from the model repository to generate corresponding activity diagrams.
[0079] First of all, the model repository 103 has the ability to automatically identify and match programming language types. The source code files in the system may use different programming languages (such as Java, Python, C++, etc.), and each language has unique syntax and structure. To ensure that the model repository can correctly process and store abstract models generated in different programming languages, the model repository uses the file extensions of the source code files to automatically identify the programming language types. For example, files with the extension ".java" correspond to the Java language, while files with the extension ".py" correspond to the Python language. In this way, the model repository can store the corresponding abstract models under the corresponding language classification according to the identified language type.
[0080] During the storage process, the model repository 103 not only stores the abstract models in a classified manner, but also improves the efficiency of data retrieval by establishing an index structure. For each programming language's abstract model, the repository generates a unique identifier (ID) for quickly locating the model. This identifier is not only related to the source code file name, but also includes the version information of the source code and other metadata, such as creation time, modification time, etc. Through these indexes, the repository can support fast query and retrieval of model data. Especially when dealing with a large number of code files, this indexing mechanism can significantly improve the response speed of the system.
[0081] In addition, the model repository 103 supports version management. In the actual software development process, the source code is continuously updated and iterated over time. To track different versions of the source code and their corresponding abstract models, the model repository can generate historical version records for each abstract model. Whenever the source code is modified, the repository stores the new abstract model as a new version and retains the records of the old versions. This version management function enables developers to easily view the state of the code at different time points or compare the changes between different versions. At the same time, if the activity diagram generator needs to generate activity diagrams of historical versions, the model repository can provide the corresponding historical abstract model data to ensure that the activity diagrams reflect the code logic of a specific version.
[0082] The model repository 103 also has data compression and optimization functions. In large projects, source code files can be very large, and the corresponding abstract models will also occupy a large amount of storage space. To optimize storage efficiency, the model repository can compress the abstract models. This compression is not a simple file compression, but an optimization for the model structure, such as merging redundant nodes, removing useless intermediate data, simplifying repetitive code segments, etc. Through these optimization measures, the model repository can minimize the use of storage space without affecting data integrity. At the same time, the compressed model data is also more efficient in retrieval and transmission, thus improving the performance of the entire system.
[0083] During the operation of the system, the model repository 103 maintains close interaction with the activity diagram generator 104. When the activity diagram generator needs to generate an activity diagram for a certain piece of code, it sends a request to the model repository. The model repository quickly retrieves the corresponding programming language abstract model according to the information in the request, such as the source code file name, programming language type, version number, etc., and passes it to the activity diagram generator for subsequent processing. The efficient retrieval mechanism of the model repository ensures that the system can quickly respond to the needs of developers and generate real-time and accurate activity diagrams.
[0084] To ensure data consistency and integrity, the model repository also has concurrent control and access permission management functions. In a multi-person collaborative development environment, multiple developers may modify the same code library simultaneously. The model repository uses a concurrent control mechanism to ensure that data does not conflict or get lost when multiple users access and modify the abstract model at the same time. At the same time, the model repository can restrict the access or modification operations of specific models according to the permissions of different users, ensuring the security and stability of the system.
[0085] Generally speaking, the model repository 103 is not just a storage module, but the core data management center of the entire system. Through a series of functions such as automatic programming language recognition, classified storage, version management, data compression optimization, and concurrent control, the model repository provides strong support for the efficient management and quick retrieval of abstract models.
[0086] Furthermore, the model repository stores the programming language abstract models through a hash table; and uses the extension name and unique identifier of the source code file as the hash key to quickly retrieve the abstract models of different programming languages.
[0087] A hash table is an efficient data structure, and its main feature is to directly map data to the storage location by calculating the hash key to achieve quick retrieval. In this system, the model repository generates a unique storage location for each programming language's abstract model through the hash table, which greatly improves the data retrieval efficiency.
[0088] The generation method of the hash key combines the extension of the source code file and the unique identifier. The extension of the source code file is used to identify the programming language type of the code. For example, ".java" corresponds to the Java language, while ".py" corresponds to the Python language. The unique identifier is used to ensure that the hash key of each file is unique, avoiding conflicts that may occur when different source code files may have the same extension. This combination ensures that the system can quickly and accurately find the corresponding data when storing and retrieving the programming language abstraction model, and can efficiently handle it whether in multiple files of the same programming language or in projects of multiple different programming languages.
[0089] Furthermore, the model repository includes a version management unit for storing and retrieving programming language abstraction models of multiple versions of the same source code file and comparing the differences between different versions.
[0090] The version management unit in the model repository undertakes the key function of version control for the programming language abstraction model. The version management unit allows the system to store programming language abstraction models of multiple versions of the same source code file and ensure that each version can be independently stored and retrieved according to the modification and iteration of the source code. In this way, when the source code changes, the system can not only save the latest version but also retain the historical versions for subsequent viewing and use.
[0091] In actual operation, when the source code file is modified, the version management unit stores the updated programming language abstraction model as a new version and generates a unique identifier for each version. These identifiers ensure that each version can be accurately traced and located. Users can retrieve the source code version at a specific time point through the version management unit and generate the corresponding activity diagram, which is convenient for backtracking and reviewing the code history.
[0092] In addition, the version management unit also has a comparison function, which can analyze and display the differences between different versions. The system identifies the added, deleted, or modified parts by comparing the programming language abstraction models of multiple versions. This function is extremely useful for developers during code maintenance, update, or debugging, because it can intuitively show the specific changes in the source code between different versions and help developers better understand and manage the evolution of the code.
[0093] The activity diagram generator 104 is connected to the model repository and is used to generate an activity diagram according to the programming language abstraction model. The activity diagram includes an activity diagram generated based on the logical code and an activity diagram generated based on the code comments.
[0094] The activity diagram generator 104 is a key component in this system. Its main function is to extract the abstract model of the programming language from the model repository 103 and generate an accurate activity diagram based on this model. The operation process of this component is complex and involves various technical details. Its work is not only to simply convert the abstract model into a graphical representation, but also to ensure the accuracy, readability, and logical consistency with the source code through multiple steps.
[0095] First, the activity diagram generator 104 extracts the required abstract model of the programming language from the model repository 103. This model is generated by the model transformation engine 102 and contains the core structural information of the source code, such as method call relationships, statement execution order, and other logical flows. The first step of the activity diagram generator is to parse this information and establish a relationship network between various logical units, especially the hierarchical and dependency relationships between function calls. This parsing process requires a detailed analysis of each node and connection in the abstract model to ensure that the generated activity diagram can reflect the actual code execution process.
[0096] One of the cores of the activity diagram generator is its ability to generate two types of activity diagrams: the activity diagram based on logical code and the activity diagram based on code comments. For the activity diagram based on logical code, the generator needs to directly extract the method call chain in the code and display information such as the call order, conditional judgment, and loop structure between various methods. This process requires the activity diagram generator to accurately identify the control flow in the code, such as `if-else` branches, `for` and `while` loops, exception handling, etc. Each control flow structure has a different representation in the activity diagram, and the activity diagram generator needs to convert them into corresponding graphical elements, such as branch nodes, loop nodes, etc., to ensure that the activity diagram accurately presents the logical relationships in the code.
[0097] For the activity diagram based on code comments, the activity diagram generator 104 focuses on presenting the business logic. The comment information is usually added by developers in the code to explain the function or business process of a certain code segment. This type of activity diagram mainly serves non-technical personnel, so the generator needs to extract the comments related to the business logic in the code, ignore the technical details, and only retain the high-level business process reflected in the comments. The generator matches the comments with the corresponding code segments and graphically displays the business logic expressed by these comments, enabling business personnel to quickly understand the business operation of the system through the activity diagram.
[0098] Another key function of the activity diagram generator 104 is the optimized generation of graphical elements. The activity diagram not only needs to accurately reflect the code logic but also be easy to understand and aesthetically pleasing. To this end, the activity diagram generator must automatically adjust the graphical layout according to the logical hierarchy in the abstract model, making the diagram clear and intuitive. The generator adjusts the size and position of each node based on its role and importance in the code execution. Important nodes (e.g., the main function or key logic control points) are usually enlarged and placed at the center of the diagram, while secondary nodes are shown in a smaller form. Through this dynamic adjustment, the activity diagram generator can generate activity diagrams with good readability, reduce the complexity of the diagram, and avoid overcrowding between nodes.
[0099] In addition, the activity diagram generator 104 also has the ability to generate interactive activity diagrams. In a modern software development environment, static activity diagrams may not fully meet the needs of developers and business personnel. To address this issue, the activity diagram generator can generate an interactive graphical interface where users can click on a node in the activity diagram to view the corresponding source code snippet or click on the connection line to view the specific call relationship or execution path. This interactive feature not only improves the usability of the activity diagram but also provides developers with a convenient means to quickly locate and understand the code.
[0100] The efficiency of the activity diagram generator is particularly important when dealing with large codebases. To this end, the generator supports parallel processing and incremental generation. For a large codebase, the activity diagram generator can allocate multiple threads or processes to handle different code modules separately and generate the activity diagram in parallel. For incremental updates to the code, the activity diagram generator can regenerate only the modified part of the code without having to completely reparse and generate the activity diagram for the entire system. This parallel and incremental generation mechanism greatly improves the performance of the system, especially when dealing with complex and frequently iterated projects, saving a significant amount of time and computing resources.
[0101] The activity diagram generator 104 can also adapt to a variety of different graphical output formats. The generator can not only generate static image files (such as PNG, SVG, etc.) but also support generating embeddable HTML pages, PDF documents, and other formats, facilitating users to output and share according to different needs. This multi-format support ensures that the activity diagram can be applied to different scenarios, whether it is for internal code review or external customer presentation, the generator can meet the requirements.
[0102] Generally speaking, the activity diagram generator 104 plays a core role in this system in transforming the abstract model into a specific visual graph. Through efficient model parsing, logically precise graph generation, graph layout optimization, interactive function support, and multi-format output, the activity diagram generator ensures that the code logic and business processes can be presented to users in the most intuitive and clear way.
[0103] Furthermore, the activity diagram generator includes a multi-perspective activity diagram generation unit, which is used to generate activity diagrams from multiple perspectives according to different user requirements. Among them, the perspectives include:
[0104] The developer perspective, which is used to display the detailed code execution logic;
[0105] The businessperson perspective, which is used to generate a simplified business logic activity diagram based on code comments;
[0106] The historical version perspective, which is used to display the function iteration or code modification history of the system.
[0107] This multi-perspective generation mechanism ensures that the activity diagram can not only meet the needs of technical developers, but also provide a simplified display of business processes for businesspersons and project managers, and even help developers trace the historical modifications of the code.
[0108] From the developer perspective, the activity diagram generated by the multi-perspective activity diagram generation unit focuses on displaying the detailed code execution logic. The activity diagram from this perspective includes all key code paths, function calls, conditional judgments, loops and other structures, helping developers understand the overall logic flow of the system and providing visual support for debugging, optimization, and function expansion. This perspective helps technical personnel quickly locate problems and understand complex code interactions.
[0109] The businessperson perspective focuses on generating a simplified business logic activity diagram. By analyzing the comment information in the code, the multi-perspective activity diagram generation unit can extract the content related to the business process and ignore the specific technical implementation details. Such activity diagrams are designed to help non-technical personnel, such as project managers, product managers, or customers, quickly understand the business logic flow of the system, facilitating communication and decision-making. This simplified graphical display makes complex technical details intuitive and clear, promoting the efficient connection between business requirements and technical implementation.
[0110] In addition, the historical version perspective allows users to view the process of system function iteration or code modification. The multi-perspective activity diagram generation unit obtains the codes of different historical versions through the version management unit and generates an activity diagram showing the functional changes. In this way, users can intuitively see the functions and code structures added, deleted, or modified between different versions, helping the development team trace problems or understand how a certain function has evolved. This function is particularly important in code review, system upgrade, or project handover.
[0111] By combining different perspectives, the activity diagram generator provides users with flexible visualization support, enabling users of different roles to quickly obtain system information relevant to their needs.
[0112] Furthermore, the activity diagram generator converts each logical node in the programming language abstract model into a visual activity diagram node through a graphics rendering engine and draws the connections between the nodes according to the function call relationships.
[0113] The activity diagram generator relies on the graphics rendering engine to convert the logical nodes in the programming language abstract model into visual activity diagram nodes. This conversion process not only involves the accurate parsing of code logic but also requires presenting the abstract program structure in a graphical way, enabling users to intuitively understand the logical flow of the system.
[0114] In this system, through the aforementioned parsing and conversion, the programming language abstract model has abstracted the basic structure of the code, function call relationships, conditional judgments, etc. into nodes and relationships in the model. The task of the graphics rendering engine is to convert these nodes into specific graphical elements on this basis, usually presented in the form of geometric shapes (such as rectangles, circles, or diamonds). Each logical node (such as a function, method, or class) will be converted into a corresponding activity diagram node, and the shape, size, and position of these nodes are dynamically calculated and drawn by the rendering engine.
[0115] In addition to the generation of nodes, the graphics rendering engine is also responsible for drawing the connections between each node according to the function call relationships. The role of the connections is to show the call and dependency relationships between each logical module in the program and the transmission of data streams. The rendering engine not only needs to consider the starting and ending points of the connections but also optimize the path of the connection lines according to the node layout to avoid excessive line crossings, thereby keeping the graph clean and legible. In some complex code structures, the rendering engine may need to use curves or broken lines to ensure that the lines can bypass other nodes without obscuring the main structure of the graph.
[0116] In addition, the graphics rendering engine can adjust the styles of nodes and connections according to user requirements. For example, developers can choose to highlight certain key function nodes, or adjust the color or size of nodes based on the frequency and importance of function calls. This flexible graphics rendering method can ensure that the generated activity diagram not only accurately shows the logical relationships of the code, but also provides a personalized graphical display according to different user concerns.
[0117] Through the graphics rendering engine, the activity diagram generator can transform the abstract programming model into an intuitive and clear activity diagram, enabling users to better understand the operating logic of the system and the interrelationships between modules.
[0118] Furthermore, the layout module includes a graphics optimization unit that adjusts the node layout of the activity diagram based on a multi-objective optimization algorithm to reduce the distance deviation between nodes, reduce the number of crossings of connection lines, and optimize the alignment degree of nodes. The multi-objective optimization algorithm is solved through the constraint optimization model provided by the following formula (1):
[0119]
[0120] where, represents the logical connection weight between the th and the th nodes, and the logical connection weight is determined based on the call relationship and dependency degree of the nodes; represents the Euclidean distance between the th and the th nodes; is the ideal distance determined according to the logical relationship; represents the length of the th connection line; is the length of the longest connection line in this diagram; represents the number of crossings of the th connection line; represents the alignment deviation of the th node; is the number of connection lines with the most crossings in this diagram; is the maximum alignment deviation in the current diagram; represents the distance between the th node and the boundary of the activity diagram; is the minimum boundary distance among all nodes; represents the number of nodes; represents the number of connection lines; , , and are trade-off parameters; It is an adjustment parameter for balancing the influence of the connection line length and the number of crossings.
[0121] The graph optimization unit adjusts the node layout in the activity graph based on a multi-objective optimization algorithm, so that the distance between nodes, the crossing of connection lines, the alignment of nodes, and the distance between nodes and the graph boundary can all reach the optimal state. To achieve this goal, the system solves through a constraint optimization model defined by formula (1).
[0122] The core of formula (1) is a multi-objective optimization function, whose goal is to simultaneously minimize multiple parameters related to the graph layout. First, the first term in the formula:
[0123] represents the balance between the logical connection and the physical distance between each pair of nodes in the activity graph. Specifically, represents the th node and the th node. The weight of the logical connection. This weight is calculated based on the call relationship and dependency degree between nodes. The higher the weight, the closer the logical relationship between the two nodes, so the system will try to place them closer in the graph.
[0124] Logical connection weight is calculated according to the call relationship and dependency degree between two nodes (i.e., the th node and the th node), reflecting the tightness of their interaction in the code logic. The higher the weight, the closer their logical relationship, so when laying out the activity graph, the system tends to place them closer. The following uses a simple example to illustrate how this weight is calculated.
[0125] Suppose there is a program with four functions: A, B, C, and D. The call relationships between these functions are as follows:
[0126] 1. A calls B and C.
[0127] 2. B calls D.
[0128] 3. C depends on the result of D.
[0129] According to this example, the calculation of the logical connection weight can be based on the following factors:
[0130] 1. Direct call relationship: If a node directly calls another node, this indicates a direct logical connection between the two. The system assigns a higher weight to this direct call relationship.
[0131] 2. Indirect call relationship: If two nodes do not directly call each other, but are indirectly connected through other nodes, their logical connection is relatively weak, but still exists. The system can assign appropriate weights to indirect calls based on the length and frequency of the call chain.
[0132] Dependency relationship: If the execution of one node depends on the result of another node, then their logical connection is also tight.
[0133] Numerical weights can be assigned to the above logical relationships:
[0134] Direct call relationship: Assume the weight of a direct call is 1.0.
[0135] Indirect call relationship: Assume the weight of an indirect call is 0.5.
[0136] Dependency relationship: Assume the weight of a dependency relationship is 0.8.
[0137] Calculate the logical connection weight between each pair of nodes based on these weights :
[0138] (A calls B, direct call);
[0139] (A calls C, direct call);
[0140] (A indirectly calls D through B);
[0141] (B calls D, direct call);
[0142] (C depends on the result of D).
[0143] Therefore, based on these weights, when generating the activity diagram, the system will tend to place A closer to B and C, B closer to D, and C also closer to D. Relatively speaking, the distance between A and D will be slightly farther because their logical connection is weaker.
[0144] In this way, the system can reasonably layout the nodes in the activity diagram according to the logical connection between each node, ensuring that nodes with close logical connections are close in the diagram, enhancing the readability and logical coherence of the activity diagram.
[0145] Denote the th node and the th node. The Euclidean distance between them calculates the physical position difference of the nodes in the two-dimensional plane.
[0146] The second term of the formula:
[0147] It involves the optimization of connection lines. represents the length of the th connection line, represents the th number of intersections of the connection line. Minimizing the connection line length can reduce visual clutter and make the graph more compact. At the same time, reducing the number of connection line intersections can prevent the lines from overlapping in the graph, thereby enhancing the neatness and readability of the activity diagram. By adjusting the trade-off parameter , the system can flexibly weigh the impact of the connection line length and the number of intersections on the layout.
[0148] The third term:
[0149] is used to focus on the alignment of nodes. represents the alignment deviation of the th node, which measures the degree of deviation of the node from the ideal alignment position. For some activity diagrams, keeping the nodes aligned horizontally or vertically helps to structure and clarify the graph. The ideal alignment position can be obtained from experimental data or set according to expert knowledge.
[0150] The fourth term:
[0151] is used to control the distance between the nodes and the graph boundary. represents the distance between the th node and the boundary of the activity diagram. By minimizing it, the system can ensure that the nodes do not get too close to the edge of the graph, avoiding the graph layout being too compact or having nodes outside the boundary. This optimization plays a key role in ensuring the overall aesthetics and visual balance of the activity diagram.
[0152] In the optimization formula, represents the number of nodes, represents the number of connection lines, , , , are trade-off parameters, and their role is to adjust the priorities of different optimization objectives. By adjusting the values of these parameters, the system can find a balance among maintaining node distances, reducing line intersections, optimizing alignment, and maintaining boundary distances. In Formula 1, , , , and have recommended values of 0.2, 0.3, 0.5, 0.6, and 0.4 respectively.
[0153] The layout module 105 is used to determine the positions, sizes, and shapes of each node in the activity diagram by parsing the method call relationships and code execution paths in the programming language abstraction model, and generate connection lines between the nodes.
[0154] The layout module 10 plays a crucial role in the system. Its task is to determine the positions, sizes, and shapes of each node in the activity diagram according to the information in the programming language abstraction model, and at the same time generate connection lines between the nodes. This process not only involves the parsing and restoration of code logic, but also requires intelligent optimization in graphical layout to ensure the readability and aesthetics of the activity diagram.
[0155] The layout module 105 first needs to parse the programming language abstraction model generated from the model transformation engine 102. This model includes detailed information on method call relationships, code execution order, and control structures (such as conditional branches and loops) in the source code. By analyzing this information, the layout module identifies the logical nodes in the activity diagram. Usually, each node represents a method, function, or control structure in the source code. During this process, the layout module must accurately capture the hierarchical structure of the code logic, such as the parent-child function call relationship, nested logic in loops, and branch structure of conditional judgments. By parsing these structures, the layout module can ensure that the placement position of each node conforms to its logical level in the code.
[0156] To ensure that the generated activity diagram has good visualization effects, the layout module 105 uses a set of complex algorithms to dynamically adjust the positions and sizes of the nodes. For a complex code structure, the number of nodes may be very large. Without a reasonable layout algorithm, the activity diagram is easily chaotic and difficult to read. Therefore, the layout module adjusts the size and position of each node according to its importance and frequency in code execution. For example, the main function or frequently called functions will be enlarged and placed at the center or prominent position of the diagram, while the secondary nodes will be reduced and placed at relatively less important positions. This dynamic adjustment of node sizes helps users quickly identify the key nodes in the activity diagram and improves the overall readability of the diagram.
[0157] In terms of node layout, the layout module 105 reduces the problem of crossing connection lines between nodes through optimization algorithms. Crossing of connection lines is a major problem that affects the readability of activity diagrams, especially when there are many nodes and complex logic. Too many crossing connection lines will make the activity diagram difficult to understand. In order to solve this problem, the layout module uses force-directed algorithms, hierarchical layout algorithms and other technologies to reasonably arrange the layout of nodes, so that logically related nodes are as close as possible, avoid long-distance connections, and minimize the crossing of connection lines. The force-directed algorithm simulates the gravitational and repulsive forces in physics to ensure that related nodes are close and unrelated nodes are separated, thereby forming a clear graph structure. The hierarchical layout algorithm is used to process code logic with a clear hierarchical structure, such as function call chains and recursive calls. These logics are usually expressed in a tree structure. The layout module arranges these nodes in layers to ensure a clear presentation of the call chain.
[0158] In addition to optimizing the position and size, the layout module 105 also needs to intelligently adjust the shape of each node. The shape of the node should not only match its logical function, but also maintain uniformity and beauty in the graph. For example, function call nodes are usually represented by rectangles, while conditional judgment nodes may be represented by diamonds or circles. Such shape distinction helps users to quickly understand the function of each node visually. The layout module automatically selects appropriate graphic elements for display by parsing the type of each node to ensure the simplicity and consistency of the graph.
[0159] The layout module 105 is also responsible for generating connecting lines between nodes, which represent method calls or control flow transfers in the code. The layout module not only needs to determine the starting and ending points of each connecting line, but also needs to optimize the path of the connecting line to ensure that the connecting line does not overlap with the node and avoids crossing with other connecting lines as much as possible. The optimization of the connecting line usually uses Bezier curves or broken lines, so that the lines can flexibly adjust the path in a limited space and maintain the logical relationship with the nodes. For some complex code logic, the connecting line may need to be bent multiple times to avoid interference from other nodes or lines. The layout module can dynamically adjust these paths according to the complexity of the graphics, so that the activity diagram as a whole looks more concise and clear.
[0160] In some advanced applications, the layout module 105 can also implement the function of dynamic layout adjustment. For very complex systems, users may want to perform interactive operations on the activity diagram, such as manually moving nodes, resizing nodes, or viewing the detailed logic of a certain part of the code. The layout module supports these interactive functions, allowing users to freely adjust the graphic layout after the activity diagram is generated, while maintaining the overall consistency of the graphic structure. To support this interaction, the layout module also provides the ability to update in real time. When the user manually modifies the graphic layout, the system will automatically recalculate the layout of other nodes and lines to ensure that the entire graphic is always in the optimal state.
[0161] The design of the layout module 105 also takes into account the adaptability to different screen sizes and output devices. Whether the activity diagram is displayed on a large screen or printed out, the layout module can automatically adjust the scaling ratio and layout structure of the graphic according to the resolution and size of the output device. In this way, no matter on which platform it is displayed, the activity diagram can maintain clear readability and adapt to different display environments.
[0162] Generally speaking, the layout module 105 is not just a simple graphic generation tool. Through complex algorithms and intelligent optimization technologies, it ensures that the generated activity diagram is visually clear and beautiful, and at the same time can accurately reflect the logical structure of the code.
[0163] Furthermore, the described automated code activity diagram generation system further includes a code processing flow complexity adjustment unit. The code processing flow complexity adjustment unit is used to calculate the complexity of the code processing flow according to the abstract syntax tree generated from the source code, and adjust the number of nodes in the activity diagram according to the code processing complexity. Among them, the calculation formula of the code processing complexity is calculated using the following formula (2):
[0164]
[0165] Among them, represents the code processing flow complexity; represents the number of execution statements in the th code snippet; represents the branch depth in the th code snippet; represents the call chain length of the th function; represents the dependency depth of the th function; represents the number of parameters of the th function; and respectively represent the number of code snippets and functions;
[0166] The code processing flow complexity adjustment unit adjusts the number of nodes and the level of detail of the activity diagram based on the following rules:
[0167] When the calculated code processing complexity exceeds a preset first threshold, the details inside the function in the activity diagram are merged into a single node, and only the high-level logic of the function call sequence is retained;
[0168] When the calculated code processing complexity exceeds a preset second threshold, multiple related code segments in the activity diagram are merged into a single node; among them, the second threshold is greater than the first threshold;
[0169] When the calculated code processing complexity is lower than a preset third threshold, the detailed execution logic inside the function is shown in the activity diagram, including branches, loops, and specific operation steps.
[0170] The described automated code activity diagram generation system further includes a code processing flow complexity adjustment unit. The function of this unit is to calculate the complexity of the processing flow based on the code structure after the source code generates an abstract syntax tree, and dynamically adjust the number of nodes and the level of detail of the activity diagram according to the calculated complexity value. The purpose of this design is to ensure that when the code structure is relatively complex, the readability of the activity diagram will not be affected by excessive details, and when the code is relatively simple, more detailed execution logic can be shown.
[0171] The code processing flow complexity adjustment unit uses formula (2) to calculate the complexity of the code processing flow, where represents the overall complexity of the code, and the specific calculation includes two parts. The first part:
[0172] is used to evaluate the complexity of the code segment, represents the number of execution statements in the th code segment, and represents the branch depth of this code segment. The branch depth usually refers to the nesting level of conditional judgments or loop structures. The deeper the nesting, the higher the branch depth. By dividing the number of execution statements by , the complexity calculation can reasonably reflect the impact of the branch structure on the complexity of the code logic.
[0173] The second part:
[0174] is used to measure the complexity of function calls. represents the call chain length of the th function, that is, how many other functions this function calls. Indicates the depth of dependence of the function, which refers to the level of function calls or the degree of nesting of dependencies. Indicates the number of parameters of the function. Generally, the more parameters a function has, the higher its complexity. Therefore, the length of the function call chain and the depth of dependence multiply the complexity.
[0175] Based on this formula, the system can dynamically adjust the number of nodes and the level of detail of the generated activity diagram. Specifically, when the code processing complexity exceeds a preset first threshold, the system determines that the code structure is relatively complex. Therefore, in the activity diagram, the details inside the function are merged into a single node, and only the high-level logic of the function call sequence is retained, which helps to simplify the graph and avoid showing too many details that make the graph overly complex. If it further exceeds a second threshold (which is greater than the first threshold), the system will take more radical measures to merge multiple related code segments into one node to reduce the number of nodes in the activity diagram and make the activity diagram more concise and clear.
[0176] Conversely, if the code processing complexity is lower than a preset third threshold, the system determines that the code structure is relatively simple. Therefore, more details will be shown in the activity diagram. In this case, the activity diagram will include the detailed execution logic inside the function, showing each branch, loop, and specific operation steps to help users deeply understand the specific process of code execution.
[0177] In the above embodiments, an automated code activity diagram generation system is provided. Correspondingly, the present application also provides an automated code activity diagram generation method. Please refer to Figure 2 , which is a flowchart of an embodiment of an automated code activity diagram generation method of the present application. Since this embodiment, that is, the second embodiment, is basically similar to the first embodiment, the description is relatively simple, and for the relevant parts, refer to the partial description of the first embodiment. The method embodiments described below are merely illustrative.
[0178] An automated code activity diagram generation method provided by the second embodiment of the present application includes:
[0179] Step S201: Extract the source code from the code file and generate an abstract syntax tree corresponding to the source code;
[0180] Step S202: Based on the query view transformation specification operation, use a set of configured parsing rules to parse the method call relationship and statement execution order in the abstract syntax tree, and convert the source code into a programming language abstract model;
[0181] Step S203: Store the programming language abstraction model in the model repository, and identify the programming language type according to the extension of the source code file to match the corresponding programming language abstraction model;
[0182] Step S204: Through the model repository, generate an activity diagram based on the programming language abstraction model, where the activity diagram includes an activity diagram generated based on logical code and an activity diagram generated based on code comments;
[0183] Step S205: By parsing the method call relationships and code execution paths in the programming language abstraction model, determine the positions, sizes, and shapes of the nodes in the activity diagram, and generate connection lines between the nodes.
[0184] Although this application is disclosed above with preferred embodiments, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application should be subject to the scope defined by the claims of this application.
Claims
1. An automatic code activity diagram generation system, characterized in that: include: A source code parsing module, used to extract source code from a code file and generate an abstract syntax tree corresponding to the source code; A model conversion engine, based on a query view conversion specification operation, is configured with a set of parsing rules for parsing the method call relationship and statement execution order in the abstract syntax tree, and converting the source code into a programming language abstract model; A model warehouse, used to store the programming language abstract model and identify the programming language type according to the extension of the source code file to match the corresponding programming language abstract model; An activity diagram generator, connected to the model warehouse, for generating an activity diagram according to the programming language abstract model, wherein the activity diagram includes an activity diagram generated based on logic code and an activity diagram generated based on code comments; A layout module, used to determine the position, size and shape of each node in the activity diagram by parsing the method call relationship and code execution path in the programming language abstract model, and generate connection lines between nodes; The code processing flow complexity adjustment unit is used to calculate the complexity of the code processing flow according to the abstract syntax tree generated by the source code, and adjust the number of nodes of the activity diagram according to the code processing complexity, wherein the calculation formula of the code processing complexity is calculated using the following formula (2): in, Indicates the complexity of the code processing flow; Indicates The number of executed statements in a code snippet; Indicates The branch depth in a code snippet; Indicates The length of the function call chain; Indicates The dependency depth of a function; Indicates The number of parameters of the function; and Represents the number of code snippets and functions respectively; The code processing flow complexity adjustment unit adjusts the number of nodes and detail of the activity diagram based on the following rules: When the calculated code processing complexity When the preset first threshold is exceeded, the details inside the function are merged into a single node in the activity diagram, and only the high-level logic of the function call sequence is retained; When the calculated code processing complexity When a preset second threshold is exceeded, multiple related code snippets are merged into a single node in the activity graph; wherein the second threshold is greater than the first threshold; When the calculated code processing complexity When the value is lower than the preset third threshold, the detailed execution logic within the function is displayed in the activity diagram, including branches, loops, and specific operation steps.
2. The automatic code activity diagram generation system according to claim 1, characterized in that: The source code parsing module performs lexical analysis on the source code through the cooperation of the lexical analyzer and the syntax analyzer, and extracts keywords, identifiers and symbols in the source code; And build an abstract syntax tree through grammatical analysis.
3. The automatic code activity diagram generation system according to claim 1, characterized in that: The source code parsing module is used to parse source code of multiple programming languages, including Java, Python and C++, and dynamically adjust the parsing process according to the grammatical rules of the programming languages.
4. The automatic code activity diagram generation system according to claim 1, characterized in that: The model warehouse stores programming language abstract models through a hash table, and uses the extension name and unique identifier of the source code file as a hash key to quickly retrieve abstract models of different programming languages.
5. The automatic code activity diagram generation system according to claim 1, characterized in that: The model repository includes a version management unit for storing and retrieving multiple versions of programming language abstract models of the same source code file and comparing the differences between different versions.
6. The automatic code activity diagram generation system according to claim 1, characterized in that: The activity diagram generator includes a multi-perspective activity diagram generating unit, and the multi-perspective activity diagram generating unit is used to generate an activity diagram from multiple perspectives according to different user requirements, wherein the perspectives include: The first perspective is used to show the detailed code execution logic; The second perspective is used to generate simplified business logic activity diagrams based on code comments; The historical version perspective is used to display the system's functional iteration or code modification history.
7. The automatic code activity diagram generation system according to claim 1, characterized in that: The activity diagram generator converts each logic node in the programming language abstract model into a visual activity diagram node through a graphic rendering engine, and draws the connection lines between the nodes according to the function call relationship.
8. The automatic code activity diagram generation system according to claim 1, characterized in that: The layout module includes a graph optimization unit, which adjusts the node layout of the activity graph based on a multi-objective optimization algorithm to reduce the distance deviation between nodes, reduce the number of crossings of connection lines, and optimize the alignment of nodes. The multi-objective optimization algorithm is solved by a constraint optimization model provided by the following formula (1): in, Indicates and The logical connection weight between nodes is determined based on the calling relationship and dependency degree of the nodes; Indicates and The Euclidean distance between nodes; is the ideal distance determined based on logical relationships; Indicates The length of the connecting line; is the length of the longest connecting line in the figure; Indicates The number of crossings of the connecting lines; Indicates The alignment deviation of nodes; is the number of lines with the most crossing times in the graph; is the maximum alignment deviation in the current graph; Indicates The distance between a node and the activity graph boundary; is the minimum boundary distance among all nodes; Indicates the number of nodes; Indicates the number of connection lines; , , and is a trade-off parameter; An adjustment parameter to balance the effect of connection line length and number of crossovers.
9. A method for generating an automated code activity diagram, characterized in that: include: Extracting source code from a code file and generating an abstract syntax tree corresponding to the source code; Based on the query view conversion specification operation, a set of configured parsing rules are used to parse the method call relationship and statement execution order in the abstract syntax tree, and convert the source code into a programming language abstract model; The programming language abstract model is stored in a model repository, and the programming language type is identified according to the extension of the source code file to match the corresponding programming language abstract model; Generate an activity diagram based on the programming language abstract model through the model warehouse, wherein the activity diagram includes an activity diagram generated based on logic code and an activity diagram generated based on code comments; By parsing the method call relationship and code execution path in the programming language abstract model, the position, size and shape of each node in the activity diagram are determined, and connecting lines between the nodes are generated; According to the abstract syntax tree generated from the source code, the complexity of the code processing flow is calculated, and the number of nodes in the activity diagram is adjusted according to the complexity of the code processing. The calculation formula of the code processing complexity is calculated using the following formula (2): in, Indicates the complexity of the code processing flow; Indicates The number of executed statements in a code snippet; Indicates The branch depth in a code snippet; Indicates The length of the function call chain; Indicates The dependency depth of a function; Indicates The number of parameters of the function; and Represents the number of code snippets and functions respectively; The code processing flow complexity adjustment unit adjusts the number of nodes and detail of the activity diagram based on the following rules: When the calculated code processing complexity When the preset first threshold is exceeded, the details inside the function are merged into a single node in the activity diagram, and only the high-level logic of the function call sequence is retained; When the calculated code processing complexity When a preset second threshold is exceeded, multiple related code snippets are merged into a single node in the activity graph; wherein the second threshold is greater than the first threshold; When the calculated code processing complexity When the value is lower than the preset third threshold, the detailed execution logic within the function is displayed in the activity diagram, including branches, loops, and specific operation steps.
Citation Information
Patent Citations
Cloud application file version management system and method based on metadata
CN117270943A