System architecture diagram generation method and device, electronic equipment and medium

By using a large language model to convert the object code of the software system into a graphical description language file and generate a system architecture diagram, the problem of low efficiency in system architecture diagram generation in the prior art is solved, and more efficient system architecture diagram generation and update are achieved.

CN120162074APending Publication Date: 2025-06-17BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220018.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, the generation efficiency of system architecture diagrams is low, which requires users to spend a lot of time and effort to understand and draw architecture diagrams, and also affects efficiency when the architecture diagram is updated.

Method used

Use a large language model to receive the target code of the software system, determine the graphic description language file corresponding to the target code, and convert it into a system architecture diagram to achieve automated generation.

Benefits of technology

It greatly reduces the time cost of manually understanding the object code and drawing the system architecture diagram, and improves the efficiency of generating and updating the system architecture diagram.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162074A_ABST
    Figure CN120162074A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a system architecture diagram generation method and device, electronic equipment and a medium. The method specifically comprises the steps that a target code of a software system is received; the target code comprises components of the software system and a calling relation and / or a data flow relation between the components; determining a graph description language file corresponding to the target code by utilizing a large language model; the graph description language file is used for describing a graph containing nodes; wherein the nodes correspond to the components in the target code; the graph description language file comprises a connection relationship between nodes; the connection relationship is used for representing a calling relationship and / or a data flow relationship between the components; converting the graph description language file into a system architecture graph; and outputting the system architecture diagram. According to the embodiment of the invention, the generation efficiency of the system architecture diagram can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a method, apparatus, electronic device, and medium for generating a system architecture diagram. Background Art

[0002] In the field of computers, software architecture is a description of the overall structure and organization of a software system, which defines the relationships, call relationships, and / or data flow relationships and behaviors between the various components of the software system. An architecture document is a document that describes and explains the software architecture in detail. As an important part of the architecture document, a system architecture diagram can visually present the structure of the software system and the relationships between the various components.

[0003] In related technologies, users often manually draw a system architecture diagram with the help of drawing tools based on their own understanding of the software system. However, this method requires users to spend a lot of time and effort to understand and draw the architecture diagram, which greatly affects the generation efficiency of the system architecture diagram. Moreover, when there is a need to update the system architecture diagram, users need to manually modify the already drawn system architecture diagram, which also affects the update efficiency of the system architecture diagram. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method, apparatus, electronic device, and medium for generating a system architecture diagram, which can greatly improve the generation efficiency of the system architecture diagram.

[0005] The specific technical solutions are as follows:

[0006] In the first aspect of the present invention, a method for generating a system architecture diagram is provided, and the method includes:

[0007] Receiving the target code of the software system; the target code includes: components of the software system, and call relationships and / or data flow relationships between the components;

[0008] Using a large language model to determine a graphical description language file corresponding to the target code; the graphical description language file is used to describe a graph including nodes; where the nodes correspond to the components in the target code; the graphical description language file includes: connection relationships between the nodes; the connection relationships are used to represent call relationships and / or data flow relationships between the components;

[0009] Converting the graphical description language file into a system architecture diagram;

[0010] Outputting the system architecture diagram.

[0011] In the second aspect of the implementation of the present invention, a device for generating a system architecture diagram is provided. The device includes:

[0012] A receiving module, configured to receive the target code of the software system; the target code includes: components of the software system, as well as the call relationship and / or data flow relationship between components;

[0013] A description language file determination module, configured to use a large language model to determine a graphical description language file corresponding to the target code; the graphical description language file is used to describe a graph including nodes; where the nodes correspond to the components in the target code; the graphical description language file includes: the connection relationship between nodes; the connection relationship is used to represent the call relationship and / or data flow relationship between components;

[0014] A text-to-diagram conversion module, configured to convert the graphical description language file into a system architecture diagram;

[0015] An output module, configured to output the system architecture diagram.

[0016] Optionally, the training process of the large language model includes:

[0017] Collecting project code samples of the software system project;

[0018] Generating a sample of the graphical description language file corresponding to the project code sample;

[0019] Using the training samples to train the large language model; the training samples include: project code samples and graphical description language file samples; the large language model is used to represent the mapping relationship between the code and the graphical description language file.

[0020] Optionally, the graphical description language file further includes: overall attribute information of the graph, node attribute information, and edge attribute information; the edge is used to connect two nodes with a connection relationship.

[0021] Optionally, the graphical description language file further includes: function category information corresponding to the nodes.

[0022] Optionally, the text-to-diagram conversion module includes:

[0023] An analysis module, configured to analyze the graphical description language file using a graphical conversion tool, and the obtained analysis results include: overall attribute information of the graph, node attribute information, and edge attribute information;

[0024] A creation and setting module, configured to create a corresponding graphical object according to the overall attribute information of the graph and set the layout information of the nodes;

[0025] A drawing module, configured to draw nodes in the graphic object according to the layout information and node attribute information, and draw edges between two nodes with a connection relationship according to the edge attribute information.

[0026] Optionally, the parsing result further includes: the function category corresponding to the node;

[0027] The creation and setting module includes:

[0028] A layout setting module, configured to set the layout information of the nodes according to the function category corresponding to the nodes; nodes with different function categories correspond to different layout information.

[0029] Optionally, the edge attribute information includes: a data flow label;

[0030] The text-graph conversion module further includes:

[0031] A text generation module, configured to generate text corresponding to the data flow label on two nodes with a connection relationship.

[0032] Optionally, the system architecture diagram is included in the architecture document; the provider of the architecture document provides an update interface for the architecture document; the device further includes:

[0033] An update module, configured to periodically obtain the latest commit hash value of the target code from the remote repository of the software system, compare the latest commit hash value with the hash value stored locally; if the two are different, obtain the updated target code from the remote repository according to the latest commit hash value; input the updated target code into the large language model to obtain an updated graphic description language file output by the large language model; convert the updated graphic description language file into an updated system architecture diagram to obtain an updated version of the system architecture diagram; call the update interface to replace the system architecture diagram in the architecture document with the updated version of the system architecture diagram.

[0034] In a third aspect of the implementation of the present invention, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0035] The memory is used to store a computer program;

[0036] The processor, when executing the program stored on the memory, implements the foregoing method steps.

[0037] In a fourth aspect of the implementation of the present invention, a computer-readable storage medium is further provided, in which instructions are stored, and when it runs on a computer, it causes the computer to execute any one of the foregoing methods.

[0038] In the fifth aspect of the implementation of the present invention, there is also provided a computer program product, including computer programs / instructions, which, when executed by a processor, implement the method described in any one of the above.

[0039] The method, device, electronic device, and medium for generating a system architecture diagram provided by the embodiments of the present invention utilize the understanding and generation capabilities of a large language model to determine a graphical description language file corresponding to the target code, and convert the above graphical description language file into a system architecture diagram.

[0040] On the one hand, the embodiments of the present invention utilize a large language model to determine a graphical description language file corresponding to the target code, greatly reducing the time cost spent by humans in understanding the target code. On the other hand, the embodiments of the present invention automatically convert the graphical description language file into a system architecture diagram, also reducing the time cost spent by humans in drawing the system architecture diagram. Since the embodiments of the present invention can save the time cost spent by humans in understanding the target code and the time cost spent by humans in drawing the system architecture diagram, therefore, the embodiments of the present invention can greatly improve the generation efficiency of the system architecture diagram. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0042] Figure 1 It is a flowchart of the steps of the method for generating a system architecture diagram according to an embodiment of the present invention;

[0043] Figure 2 It is a schematic diagram of a graphical description language file according to an embodiment of the present invention;

[0044] Figure 3 It is a schematic diagram of a system architecture diagram according to an embodiment of the present invention;

[0045] Figure 4 It is a schematic diagram of a system architecture diagram according to an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the structure of a device for generating a system architecture diagram according to an embodiment of the present invention;

[0047] Figure 6 It is a block diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following will describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.

[0049] The method for generating the system architecture diagram according to the embodiments of the present invention can be used to automatically generate a system architecture diagram based on the target code of a software system, so as to improve the generation efficiency of the system architecture diagram.

[0050] The embodiments of the present invention do not limit the specific application scenarios corresponding to the software system. For example, the application scenarios of the software system may include: Internet scenarios, Internet of Things scenarios, or financial scenarios, etc.

[0051] The method for generating the system architecture diagram in the related art specifically includes: the user manually draws the system architecture diagram with the help of a drawing tool based on their own understanding of the software system. However, this method requires the user to spend a lot of time and effort to understand and draw the architecture diagram, which greatly affects the generation efficiency of the system architecture diagram.

[0052] In view of the technical problem of the low generation efficiency of the system architecture diagram in the related art, the embodiments of the present invention provide a method for generating a system architecture diagram, which specifically includes: receiving the target code of the software system; the above target code includes: components of the software system, and the call relationship and / or data flow relationship between the components; using a large language model to determine the graphical description language file corresponding to the above target code; the above graphical description language file is used to describe a graph including nodes; wherein, the nodes correspond to the components in the target code; the above graphical description language file includes: the connection relationship between nodes; the above connection relationship is used to represent the call relationship and / or data flow relationship between components; converting the above graphical description language file into a system architecture diagram; and outputting the above system architecture diagram.

[0053] The embodiments of the present invention utilize the understanding ability and generation ability of the large language model to determine the graphical description language file corresponding to the target code, and convert the above graphical description language file into a system architecture diagram.

[0054] On the one hand, the embodiments of the present invention use a large language model to determine the graphical description language file corresponding to the target code, which greatly reduces the time cost spent by humans in understanding the target code. On the other hand, the embodiments of the present invention automatically convert the graphical description language file into a system architecture diagram, which also reduces the time cost spent by humans in drawing the system architecture diagram. Since the embodiments of the present invention can save the time cost spent by humans in understanding the target code and the time cost spent by humans in drawing the system architecture diagram, therefore, the embodiments of the present invention can greatly improve the generation efficiency of the system architecture diagram.

[0055] The embodiments of the present invention will be described below through specific embodiments.

[0056] Refer to Figure 1 , which shows the flowchart of the steps of the method for generating the system architecture diagram according to an embodiment of the present invention. The method may specifically include the following steps:

[0057] Step 101: Receive the target code of the software system; the above target code specifically includes: components of the software system, as well as the call relationship and / or data flow relationship between components;

[0058] Step 102: Use a large language model to determine the graphical description language file corresponding to the above target code; the above graphical description language file is used to describe a graph containing nodes; among them, the nodes correspond to the components in the target code; the above graphical description language file includes: the connection relationship between nodes; the above connection relationship is used to represent the call relationship and / or data flow relationship between components;

[0059] Step 103: Convert the above graphical description language file into a system architecture diagram;

[0060] Step 104: Output the above system architecture diagram.

[0061] Figure 1 The method embodiment shown can be used to automatically generate a system architecture diagram according to the target code of the software system, so as to improve the generation efficiency of the system architecture diagram.

[0062] In step 101, a code upload interface can be provided. The above code upload interface is used to receive the target code uploaded by the user. Among them, the target code can be the initial code of the software system or the updated code of the software system.

[0063] Among them, in the case where the target code is the initial code of the software system, the embodiment of the present invention can automatically generate an initial version of the system architecture diagram according to the initial code of the software system. In the case where the target code is the updated code of the software system, the embodiment of the present invention can automatically generate an updated version of the system architecture diagram according to the updated code of the software system.

[0064] It should be noted that the system architecture diagram of the embodiment of the present invention can be included in the architecture document. The provider of the architecture document can provide an update interface for the architecture document. In the case where the target code of the software system is updated, the embodiment of the present invention can automatically generate an updated version of the system architecture diagram according to the updated code, and call the above update interface to automatically update the system architecture diagram in the architecture document.

[0065] For example, software system A corresponds to architecture document A, and the architecture document has a version number A.i. In the case where the object code of software system A is updated, the embodiments of the present invention can automatically generate an updated version of the system architecture diagram according to the updated code of the software system, and automatically update the system architecture diagram in the architecture document to the updated version of the system architecture diagram according to the above-mentioned update interface, so as to update the version number of the architecture document from version number A.i to version number A.i+1. Wherein, i can be a positive integer.

[0066] Among them, the embodiments of the present invention can regularly obtain the latest commit hash value of the object code from the remote repository of the software system, and compare the latest commit hash value with the hash value stored locally; if the two are different, it indicates that there is a new commit record in the remote repository, which means that the object code has been updated. At this time, the updated object code can be obtained from the remote repository according to the latest commit hash value, so as to realize the update of the object code of the software system.

[0067] Among them, the locally stored hash value refers to the hash value corresponding to the software version used by the current system architecture diagram. The latest commit hash value is the hash value corresponding to the latest commit version of the object code in the remote repository.

[0068] The hash value is the unique identifier of the software version, and different hash values correspond to different software versions. In practical applications, the unique hash value of the software version can be generated by using a set hash algorithm based on metadata such as the code content, committer, commit time, and commit description when the software version is committed. The set hash algorithm can include: SHA (Secure Hash Algorithm), etc.

[0069] Input the updated object code into the large language model to obtain an updated graphical description language file output by the large language model; convert the updated graphical description language file into an updated system architecture diagram to obtain an updated version of the system architecture diagram.

[0070] In step 102, the understanding ability and generation ability of the large language model can be used to determine the graphical description language file corresponding to the above-mentioned object code.

[0071] The large language model in the embodiments of the present invention is an artificial intelligence model designed to understand and generate human language. By training on a large amount of text data, the large language model can perform a wide range of tasks, including the task of generating graphical description language files in the embodiments of the present invention, and so on. The large language model is usually based on a deep learning architecture, such as the Transformer structure. The embodiments of the present application do not limit the specific type of the large language model. For example, the types of the large language model can include: GPT (Generative Pre-Trained Transformer) model and GLM (General Language Model), etc.

[0072] For example, the GPT model is a large-scale natural language generation model based on the Transformer structure. It uses a vast amount of text data for pre-training and can generate high-quality natural language texts, including articles, dialogues, abstracts, etc. The core of the GPT model is the Transformer structure, which adopts an attention mechanism, enabling the GPT model to consider all positions in the input sequence simultaneously, thereby better capturing long-range dependencies. After pre-training the GPT model, fine-tuning it can adapt to various downstream text generation tasks.

[0073] The embodiments of the present invention can perform fine-tuning training on the large language model so that the large language model can represent the mapping relationship between the code and the graphical description language file. In other words, the fine-tuning training in the embodiments of the present invention enables the large language model to have the ability to output a graphical description language file according to the input code.

[0074] Correspondingly, the training process of the large language model specifically includes:

[0075] Step A1, collect project code samples of a software system project;

[0076] Step A2, generate graphical description language file samples corresponding to the project code samples;

[0077] Step A3, input the project code samples into the large language model to obtain the prediction result of the graphical description language file output by the large language model; the large language model is used to represent the mapping relationship between the code and the graphical description language file;

[0078] Step A4, determine the loss information according to the prediction result and the graphical description language file samples;

[0079] Step A5, update the model parameters of the large language model according to the loss information.

[0080] The training from step A1 to step A5 can be fine-tuning training for the large language model.

[0081] In step A1, the components of the software system can include: the main program component and the middleware.

[0082] The main program component is usually the part of the software system that implements the core business logic. It is responsible for processing the main requests and business processes of users, coordinating the operation of each part, and is a key component of the entire software system.

[0083] For example, in an e-commerce platform, the main program component may include: function modules such as user login and registration, product display, shopping cart management, and order processing. These function modules directly face users and process various operations and business requirements of users.

[0084] Middleware is a software layer located between the operating system and the application program. It provides various general services and functions for the software system to simplify the development process, improve the maintainability and scalability of the system.

[0085] Common middleware includes: database management systems (such as MySQL), cache systems (such as Redis), message queues (such as RocketMQ), search engines (such as Elasticsearch), etc. These middleware provide functions such as data storage, cache acceleration, asynchronous communication, and fast search respectively.

[0086] The project code sample of the embodiment of the present invention specifically includes: the code of the main program component, the code of the middleware, and the code for the main program component and the middleware to work together.

[0087] In step A2, a graphic description language file sample corresponding to the project code sample can be generated by manual means or automated tools. It can be understood that the embodiment of the present invention does not limit the specific generation method of the graphic description language file sample.

[0088] The graphic description language of the embodiment of the present invention can include: DOT (Description Output Tool), SVG (Scalable Vector Graphics), etc. Among them, DOT is used to describe directed graphs and undirected graphs. Its syntax is concise and intuitive, and it can define the attributes of nodes, edges, and graphics, etc. It is widely used in drawing various types of graphics, including flowcharts, network topology diagrams, software architecture diagrams, etc.

[0089] In step A3, the training samples specifically include: the project code sample and the graphic description language file sample, where the graphic description language file sample can be used as the label value of the project code sample. The training process of the large language model can include: forward propagation and backward propagation.

[0090] Among them, forward propagation can calculate the prediction result step by step according to the model parameters of the large language model in the order from the input layer to the output layer.

[0091] Backward propagation can calculate and update the model parameters of the large language model step by step according to the loss information in the order from the output layer to the input layer. The large language model usually adopts the structure of a neural network. The model parameters of the large language model can include parameters such as the weights of the neural network. Among them, during the backward propagation process, the gradient information of the model parameters of the large language model can be determined, and this gradient information can be used to update the model parameters of the large language model. For example, backward propagation can calculate and store the gradient information of the model parameters of the large language model step by step according to the chain rule in calculus in the order from the output layer to the input layer.

[0092] During the forward propagation process of the large language model, the project code sample can be input into the large language model, and the large language model can perform forward operations on the project code sample and output the prediction result. And the loss information can be determined according to the prediction result and the graphical description language file sample.

[0093] In a specific implementation, iterative training can be performed according to multiple batches of training samples. The convergence condition of the above iteration can be: the loss information meets the preset condition. The preset condition can be: the absolute value of the difference between the loss information and the target value is less than the difference threshold, or the number of iterations exceeds the number threshold, etc. In other words, when the loss information meets the preset condition, the iteration can be ended; in this case, the target parameter value of the model parameters of the large language model can be obtained. Subsequently, the graphical description language file corresponding to the target code can be determined according to the target parameter value of the model parameters of the large language model.

[0094] The large language model usually has a large number of model parameters and is prone to overfitting on the training samples, that is, the large language model performs well on the training samples but its performance deteriorates on new and unseen data.

[0095] The loss function in the embodiment of the present invention can include a regularization term. The regularization term can effectively prevent this overfitting, enable the large language model to have better generalization ability, and be able to perform well on different data sets. The loss function is used to calculate the loss information.

[0096] For example, the L2 regularization term can make the model parameters of the large language model tend to smaller values, avoiding the over-sensitivity of the large language model to the noise in the training data due to some overly large model parameters. The L2 regularization term is half of the sum of the squares of the parameters multiplied by the regularization coefficient. In other words, square each model parameter of the large language model, then calculate the sum of the squares of all model parameters, and finally multiply by half of the regularization coefficient to obtain the L2 regularization term.

[0097] The L1 regularization term may make some model parameters become zero, achieving the effect of feature selection and reducing the complexity of the large language model. In the embodiments of the present invention, the absolute values of all model parameters of the large language model can be added together, and the resulting sum of absolute values can represent the overall deviation of the entire model parameter set from zero; multiplying the sum of absolute values by the regularization coefficient gives the L1 regularization term.

[0098] After completing the training process of the large language model, the embodiments of the present invention can also use a test set to evaluate the performance of the large language model. The performance of the large language model specifically includes: the accuracy, readability, and practicality of the generated results, etc. The embodiments of the present invention can use manual evaluation methods and automatic evaluation methods to evaluate the performance of the large language model.

[0099] Moreover, the evaluation results of the performance of the large language model specifically include: evaluation passed or evaluation failed. In the case of evaluation passed, the large language model can be deployed. In the case of evaluation failed, the model structure and hyperparameters of the large language model can be adjusted, and the adjusted large language model can be retrained and re-evaluated until the evaluation result is evaluation passed. In machine learning, a hyperparameter is a parameter set before the start of the learning process, rather than a parameter obtained through training. Hyperparameters specifically include: the number of network layers, learning rate, training batch size, or number of iterations, etc. It can be understood that the embodiments of the present invention do not limit specific hyperparameters.

[0100] The deployment of the large language model can deploy the large language model with evaluation passed to a preset hardware environment, so that the large language model can determine the graphical description language file corresponding to the target code in the preset hardware environment. The preset hardware environment can be a server environment or a client environment. Or, the preset hardware environment can be a server environment or a mobile terminal environment. It can be understood that the embodiments of the present invention do not limit specific preset hardware environments.

[0101] In practical applications, the target code can be input into a large language model, and a graphical description language file output by the large language model can be received. The graphical description language file can be a DOT script file. It can be understood that the embodiments of the present invention do not limit the specific graphical description language file.

[0102] The above-mentioned graphical description language file is used to describe a graph containing nodes; wherein, the nodes correspond to components in the target code; the above-mentioned graphical description language file includes: the connection relationship between nodes; the above-mentioned connection relationship is used to represent the call relationship and / or data flow relationship between components.

[0103] Referring to Figure 2 , a schematic diagram of a graphical description language file according to an embodiment of the present invention is shown, which includes multiple call relationship paths between components in a software system.

[0104] The first call relationship path: The user accesses component 1, component 1 calls component 3, component 3 calls component 4, component 4 calls component 3, component 3 calls component 5, and component 5 calls component 6.

[0105] The second call relationship path: The user accesses component 2, component 2 calls component 3, component 3 calls component 4, component 4 calls component 3, component 3 calls component 5, and component 5 calls component 6.

[0106] The third call relationship path: The user accesses component 1, component 1 calls component 3, component 3 calls component 5, component 5 calls component 6, component 6 calls component 7, and component 7 calls component 6.

[0107] The fourth call relationship path: The user accesses component 2, component 2 calls component 3, component 3 calls component 5, component 5 calls component 6, component 6 calls component 7, and component 7 calls component 6.

[0108] In one example, the target code corresponds to a Spring Boot project that uses mysql and ElasticSearch databases for read-write separation, uses mysql for write operations, uses ElasticSearch for query operations, and mysql automatically synchronizes data to ElasticSearch. This project uses redis as a cache middleware, provides an internal call interface InnerAPI for other services, and provides a data display interface UserAPI for external users.

[0109] For the graphical description language file generated by the embodiments of the present invention for the target code, a directed graph G can be represented by Digraph G{}, and the information of the directed graph G is as follows:

[0110]

[0111] Among them, the nodes of the directed graph G include: ElasticSearch (Elastic Search), SpringBoot (Spring Scaffold), Mysql (Relational Database), Redis (Remote Dictionary Server), UserAPI (User Application Programming Interface), and InnerAPI (Internal Interface), etc. SpringBoot can correspond to the main program component in the target code. ElasticSearch, Mysql, Redis, UserAPI, and InnerAPI can correspond to the middleware in the target code.

[0112] The connection relationships between nodes are as follows:

[0113] ElasticSearch->SpringBoot[label=”query data”] means that there is an edge from the ElasticSearch node to the SpringBoot node, and the data flow label on the edge is “query data”, that is, ElasticSearch queries data from SpringBoot.

[0114] SpringBoot->Mysql[label=”write data”] means that SpringBoot writes data to Mysql.

[0115] Mysql->ElasticSearch[label=”synchronization”] means that Mysql synchronizes data to ElasticSearch.

[0116] Redis->SpringBoot[label=”cache”] means that Redis provides caching for SpringBoot.

[0117] SpringBoot->UserAPI[label=”display”] means that SpringBoot provides data to UserAPI for display.

[0118] SpringBoot->InnerAPI[label=”invoke”] means that SpringBoot invokes InnerAPI.

[0119] The graphical description language file of the embodiment of the present invention may further include: graphical style information such as overall attribute information of the graph, node attribute information, and edge attribute information; the edge is used to connect two nodes with a connection relationship.

[0120] Among them, the overall attribute information is used to characterize the global visual features of the graph, and the overall attribute information may include: the background color of the graph, the font type, and the font size.

[0121] The node attribute information is used to characterize the visual features of the nodes in the graph. The node attribute information may include: the shape of the node, the filling style and filling color of the node, the font type and font size of the node, etc. The edge attribute information may include: the color of the edge, and the font type and font size of the data flow label on the edge, etc.

[0122] It can be understood that those skilled in the art can determine the graphical style information according to actual application requirements, and set the graphical style information in the large language model, and the large language model carries the above-mentioned graphical style information in the output graphical description language file.

[0123] For example, the attribute information of the directed graph G is as follows:

[0124] Overall attribute information of the graph:

[0125] Graph[bgcolor=“lightyellow”,fontname=“Arial”,fontsize=10]. Among them, bgcolor=“lightyellow” sets the background color of the graph to light yellow. fontname=“Arial” specifies that the font used in the graph is Arial (Song typeface). fontsize=10 sets the font size in the graph to size 10.

[0126] Node attribute information:

[0127] Node[shape=“box”,style=“filled”,fillcolor=“lightblue”,fontname=“Arial”,fontsize=10]. Among them, shape=“box” is used to set the shape of the node to a rectangle. style=“filled” is used to indicate that the filling style of the node is solid filling. fillcolor=“lightblue” is used to set the filling color of the node to light blue. fontname=“Arial” and fontsize=10 are consistent with the font attributes in the overall graph settings, so that the data flow labels on the nodes also use the Arial font and the size is 10.

[0128] Edge attribute information:

[0129] Edge[color = "darkgray", fontname = "Arial", fontsize = 9]. Among them, color = "darkgray" sets the color of the edge to dark gray. fontname = "Arial" and fontsize = 9 specify that the font of the data flow label on the edge is Arial and the size is 9 points.

[0130] In an alternative implementation, the graphical description language file may further include: function category information corresponding to the nodes. The function category information is a descriptive information used to classify the nodes in the system architecture diagram, which characterizes the architectural level where the nodes are located in the software system.

[0131] For example, the function category information corresponding to the nodes may include: presentation layer, logic layer, storage layer, etc. Classifying the nodes into function categories such as the presentation layer, logic layer, and storage layer can intuitively reflect the architectural level of the system. For example, the presentation layer may include interface components that directly interact with users, such as web front-ends or mobile application interfaces; the logic layer contains modules that process business logic and are responsible for data processing and operations; the storage layer is responsible for persistent data storage, such as a database system. This division helps developers and architects better understand the structure of the system and the relationships between its various parts, facilitating system design, maintenance, and optimization.

[0132] In step 103, a graphical conversion tool can be used to convert the graphical description language file into a system architecture diagram. Examples of graphical conversion tools may include: Graphviz (Graph Visualization Software), or PlantUML (Plant Unified Modeling Language), etc.

[0133] Specifically, the process of converting the above-mentioned graphical description language file into a system architecture diagram specifically includes:

[0134] Step B1: Parse the graphical description language file using a graphical conversion tool, and the obtained parsing results include: overall attribute information of the graph, node attribute information, and edge attribute information;

[0135] Step B2: Create a corresponding graphical object according to the overall attribute information of the graph and set the layout information of the nodes;

[0136] Step B3: Draw nodes in the graphical object according to the layout information and node attribute information, and draw edges between two nodes with a connection relationship according to the edge attribute information.

[0137] In step B1, the graph conversion tool can read the graph description language file and understand the syntax and structure in the graph description language file. These tools usually have a parser that can convert the text description in the graph description language file into an internal data structure for subsequent processing. For example, the parser will identify elements such as nodes and edges in the DOT script file and store the elements in the corresponding data structures.

[0138] In step B2, the graph object can be a canvas, a graphic container, or a specific data structure for storing and drawing each element of the graph.

[0139] The graph object is an entity that carries the overall attribute information of the graph. And the overall attribute information is a description of the global visual characteristics of the graph object.

[0140] The process of creating the corresponding graph object may include: creating a graph object according to the function for creating a graph object in the graph library, and setting the attributes of the graph object according to the overall attribute information of the graph.

[0141] If the overall attribute information indicates that the background color of the graph is light blue, then after creating the graph object, the background color attribute of the graph object will be set to light blue through the corresponding setting function. For the font type, if the overall attribute information specifies the regular script, when there are text elements in the graph object, in the operation of setting the text-related attributes, the regular script will be found from the font selection list and applied so that the text in the graph object is displayed in the regular script. As for the font size, if the overall attribute information gives a font size of 14, similarly at the text attribute setting, the attribute value of the font size will be adjusted to 14.

[0142] The layout information determines the position and arrangement of the nodes in the graph. The graph conversion tool can automatically calculate a suitable layout according to the type and attributes of the graph, as well as the relationship between the nodes and edges.

[0143] In one implementation, a force-directed layout algorithm can be used to determine the layout information of the nodes. The force-directed layout algorithm simulates the forces in a physical system, regarding the nodes as mass points and the edges as springs. There is a repulsive force between the nodes, so that they do not overlap; there is a tensile force in the edges, so that the nodes with connection relationships approach each other.

[0144] In another implementation, the parsing result may further include: the function category corresponding to the node; then the process of setting the layout information of the node may specifically include: setting the layout information of the node according to the function category corresponding to the node; nodes of different function categories correspond to different layout information.

[0145] When the parsing result includes the function categories corresponding to the nodes, different layout information can be set for the nodes according to different function categories. This way can better reflect the hierarchical structure and function division of the system, making the graph clearer and easier to understand. For example, nodes of different function categories can be placed in different areas, or different colors and shapes can be used to distinguish nodes of different function categories.

[0146] In an optional implementation, the edge attribute information may specifically include: a data flow label; then the process of converting the graphical description language file into a system architecture diagram may further include: generating a data flow label on two nodes with a connection relationship. In the system architecture diagram, the data flow label can be near the two nodes and can be presented in text form.

[0147] In the system architecture diagram, the data flow label is used to show the flow direction of data between different nodes and the business attributes carried by the data during the flow process. It can help users quickly understand the data call relationship and / or data flow relationship between various components in the system, which is crucial for analyzing the functions and performance of the system. In the embodiments of the present invention, adding a data flow label to the edge connecting two nodes can make the graph more intuitive and easier to understand.

[0148] In a software system, the data call relationship refers to the behavior of data interaction between components through interfaces, methods or services, usually presenting a two-way request-response mode, commonly seen in scenarios such as function calls, and the call information involved usually includes call methods, parameters, interface names, etc.

[0149] In a software system, the data flow relationship refers to the flow path of data from generation to processing, storage or output, generally a one-way flow from one component to another component.

[0150] It should be noted that developers can add comments in the target code, and the comments are used to explain the data flow labels involved in specific parts of the code. For example, for a function call, it can be noted in the comment that "this function processes user data input, and the data flow label is 'data flow name'". The large language model can extract the data flow label by analyzing the comments in the code. The data flow label can not only include: the data flow name, but also include: the data flow direction. For example, for "component A reads data from component B", the data flow label can be: "data flow name: the data reading flow between component A and component B, data flow direction: component A reads data from component B". For "component C writes data to component D", the data flow label can be: "data flow name: the data writing flow between component C and component D, data flow direction: component C writes data to component D".

[0151] Refer to Figure 3, showing a schematic diagram of the system architecture of an embodiment of the present invention, where the background color of the system architecture diagram can be determined according to the overall attribute information of the graph.

[0152] Figure 3 In it, the nodes of the system architecture diagram specifically include: ElasticSearch, SpringBoot, Mysql, Redis, UserAPI, and InnerAPI.

[0153] Among them, there is an edge from ElasticSearch to SpringBoot, and the data flow label on this edge is "query data", indicating that ElasticSearch provides query data to SpringBoot.

[0154] There is an edge from SpringBoot to the relational database, and the data flow label on this edge is "write data", indicating that SpringBoot writes data to the relational database.

[0155] There is an edge from the relational database to ElasticSearch, and the data flow label on this edge is "synchronize", indicating that the relational database synchronizes data to ElasticSearch.

[0156] There is an edge from the relational database to SpringBoot, and the data flow label on this edge is "cache", indicating that the relational database provides cache to SpringBoot.

[0157] There is an edge from SpringBoot to the user interface, and the data flow label on this edge is "display", indicating that SpringBoot provides data for display to the user interface.

[0158] There is an edge from SpringBoot to the internal interface, and the data flow label on this edge is "invoke", indicating that SpringBoot sends an invocation request to the internal interface.

[0159] Refer to Figure 4 , showing a schematic diagram of the system architecture of an embodiment of the present invention. Relative to Figure 3 the shown system architecture diagram, Figure 4 the shown system architecture diagram determines the layout information of the nodes according to the functional categories corresponding to the nodes. Figure 4 In it, the nodes of the presentation layer specifically include: the user interface and the internal interface. The nodes of the logic layer specifically include: SpringBoot. The nodes of the storage layer specifically include: ElasticSearch, the relational database, and the remote dictionary service. Figure 4 In the system architecture diagram, corresponding layer areas are set for the presentation layer, the logic layer, and the storage layer respectively, and the layer names and nodes are drawn in the layer areas.

[0160] Since Figure 4 the connection relationships between the nodes inFigure 3 The connection relationships among intermediate nodes are similar, so they will not be elaborated here and can be referred to each other.

[0161] In step 104, the system architecture diagram can be output to a graphic file. The system architecture diagram or the graphic file corresponding to the system architecture diagram can also be displayed on the screen for the user to view and download.

[0162] The embodiments of the present invention do not limit the specific format of the graphic file. For example, the format of the graphic file can include: PNG (Portable Network Graphics) format, JPEG (Joint Photographic Experts Group) format, etc.

[0163] In practical applications, the system architecture diagram can be included in the architecture document, and the provider of the architecture document provides an update interface for the architecture document; the method can also include: regularly obtaining the latest commit hash value of the target code from the remote repository of the software system, comparing the latest commit hash value with the hash value stored locally; if the two are different, obtaining the updated target code from the remote repository according to the latest commit hash value; inputting the updated target code into the large language model to obtain an updated graphic description language file output by the large language model; converting the updated graphic description language file into an updated system architecture diagram to obtain an updated version of the system architecture diagram; calling the update interface and using the updated version of the system architecture diagram to replace the system architecture diagram in the architecture document.

[0164] With the continuous development and evolution of the software system, the target code will be updated frequently. The embodiments of the present invention automatically generate an updated version of the system architecture diagram and automatically update the architecture document, which can achieve the consistency between the architecture document and the actual software system. This enables team members to obtain the latest system architecture diagram when viewing the architecture document, avoiding misunderstandings and incorrect decisions caused by the inconsistency between the architecture document and the actual software system.

[0165] In summary, the method for generating a system architecture diagram according to the embodiments of the present invention utilizes the understanding and generation capabilities of the large language model to determine the graphic description language file corresponding to the target code, and converts the above graphic description language file into a system architecture diagram.

[0166] On the one hand, the embodiments of the present invention utilize a large language model to determine the graphical description language file corresponding to the target code, greatly reducing the time cost of manually understanding the target code. On the other hand, the embodiments of the present invention automatically convert the graphical description language file into a system architecture diagram, also reducing the time cost of manually drawing the system architecture diagram. Since the embodiments of the present invention can save the time cost of manually understanding the target code and the time cost of manually drawing the system architecture diagram, therefore, the embodiments of the present invention can greatly improve the generation efficiency of the system architecture diagram.

[0167] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0168] Referring to Figure 5 , a structural schematic diagram of a device for generating a system architecture diagram according to an embodiment of the present invention is shown. The device for generating a system architecture diagram may specifically include the following modules: a receiving module 501, a description language file determination module 502, a text-graph conversion module 503, and an output module 504.

[0169] The receiving module 501 is configured to receive the target code of the software system; the target code includes: components of the software system, and the call relationship and / or data flow relationship between the components;

[0170] The description language file determination module 502 is configured to utilize a large language model to determine the graphical description language file corresponding to the target code; the graphical description language file is used to describe a graph including nodes; wherein, the nodes correspond to the components in the target code; the graphical description language file includes: the connection relationship between the nodes; the connection relationship is used to represent the call relationship and / or data flow relationship between the components;

[0171] The text-graph conversion module 503 is configured to convert the graphical description language file into a system architecture diagram;

[0172] The output module 504 is configured to output the system architecture diagram.

[0173] Optionally, the training process of the large language model includes:

[0174] Collect project code samples of software system projects;

[0175] Generate a graphical description language file sample corresponding to the project code sample;

[0176] Use the training samples to train the large language model; the training samples include: project code samples and graphical description language file samples; the large language model is used to represent the mapping relationship between the code and the graphical description language file.

[0177] Optionally, the graphical description language file further includes: overall attribute information of the graph, node attribute information, and edge attribute information; the edge is used to connect two nodes with a connection relationship.

[0178] Optionally, the graphical description language file further includes: function category information corresponding to the node.

[0179] Optionally, the text-to-graph conversion module 503 may specifically include:

[0180] A parsing module, configured to parse the graphical description language file using a graphical conversion tool, and the obtained parsing result includes: overall attribute information of the graph, node attribute information, and edge attribute information;

[0181] A creation and setting module, configured to create a corresponding graphical object according to the overall attribute information of the graph and set the layout information of the nodes;

[0182] A drawing module, configured to draw nodes in the graphical object according to the layout information and node attribute information, and draw an edge between two nodes with a connection relationship according to the edge attribute information.

[0183] Optionally, the parsing result further includes: the function category corresponding to the node;

[0184] The creation and setting module includes:

[0185] A layout setting module, configured to set the layout information of the nodes according to the function category corresponding to the nodes; nodes of different function categories correspond to different layout information.

[0186] Optionally, the edge attribute information includes: a data flow label;

[0187] The text-to-graph conversion module 503 may further include:

[0188] A text generation module, configured to generate text corresponding to the data flow label on two nodes with a connection relationship.

[0189] Optionally, the system architecture diagram is included in the architecture document; the provider of the architecture document provides an update interface for the architecture document; the apparatus further includes:

[0190] An update module is used to periodically obtain the latest commit hash value of the target code from the remote repository of the software system, and compare the latest commit hash value with the hash value stored locally; if the two are different, then according to the latest commit hash value, obtain the updated target code from the remote repository; input the updated target code into the large language model to obtain the updated graphical description language file output by the large language model; convert the updated graphical description language file into an updated system architecture diagram to obtain an updated version of the system architecture diagram; call the update interface and use the updated version of the system architecture diagram to replace the system architecture diagram in the architecture document.

[0191] An embodiment of the present invention also provides an electronic device, which can implement the functions of the foregoing storage node, scheduling node, ledger node, or management node.

[0192] As Figure 6 shown, the electronic device may include a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004.

[0193] The memory 1003 is used to store computer programs.

[0194] When the processor 1001 is used to execute the program stored on the memory 1003, the following steps are implemented:

[0195] Store the storage object.

[0196] Send the storage object to the coupled second storage node so that the second storage node stores and / or forwards the storage object.

[0197] After successfully storing the storage object, send the object information of the storage object to the ledger node so that the ledger node saves the ledger information of the first storage node; the ledger information includes: the object information stored by the corresponding storage node.

[0198] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0199] The communication interface is used for communication between the above electronic device and other devices.

[0200] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one generating device of a system architecture diagram located far from the aforementioned processor.

[0201] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0202] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, the computer is made to execute the method for generating a system architecture diagram described in any one of the above embodiments.

[0203] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, the computer is made to execute the method for generating a system architecture diagram described in any one of the above embodiments.

[0204] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state disk (SSD)). Examples of optical media can include DVDs (Digital Video Discs).

[0205] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0206] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0207] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A method for generating a system architecture diagram, characterized in that: The method comprises: Receive the target code of the software system; the target code includes: components of the software system, and calling relationships and / or data flow relationships between components; Determine a graph description language file corresponding to the target code by using a large language model; the graph description language file is used to describe a graph containing nodes; wherein the nodes correspond to components in the target code; the graph description language file includes: a connection relationship between nodes; the connection relationship is used to characterize a calling relationship and / or a data flow relationship between components; Converting the graphic description language file into a system architecture diagram; The system architecture diagram is output.

2. The method according to claim 1, characterized in that The training process of the large language model includes: Collecting project code samples of software system projects; the project code samples include: codes of main program components, codes of middleware, and codes of cooperation between the main program components and the middleware; Generate a graphics description language file sample corresponding to the project code sample; Inputting the project code sample into the large language model to obtain the prediction result of the graphic description language file output by the large language model; the large language model is used to characterize the mapping relationship between the code and the graphic description language file; Determining loss information according to the prediction result and the graphic description language file sample; The model parameters of the large language model are updated according to the loss information.

3. The method according to claim 1, characterized in that The graph description language file also includes: overall attribute information of the graph, node attribute information, and edge attribute information; the edge is used to connect two nodes with a connection relationship; wherein the overall attribute information is used to characterize the global visual features of the graph; and the node attribute information is used to characterize the visual features of the nodes in the graph.

4. The method according to claim 3, characterized in that The graphic description language file also includes: function category information corresponding to the node; the function category information is a descriptive information used to classify the nodes in the system architecture diagram, and is used to characterize the architecture level at which the node is located in the software system.

5. The method according to any one of claims 1 to 4, characterized in that: The converting the graphic description language file into a system architecture diagram comprises: The graph description language file is parsed using a graph conversion tool, and the parsing results obtained include: overall attribute information of the graph, node attribute information and edge attribute information; According to the overall attribute information of the graph, a corresponding graph object is created and the layout information of the nodes is set; Nodes are drawn in the graphic object according to the layout information and the node attribute information, and edges are drawn between two nodes having a connection relationship according to the edge attribute information.

6. The method according to claim 5, characterized in that The analysis result also includes: the function category corresponding to the node; The setting node layout information includes: The layout information of the node is set according to the functional category corresponding to the node; nodes of different functional categories correspond to different layout information.

7. The method according to claim 5, characterized in that The edge attribute information includes: a data flow label; the data flow label is used to show the flow direction of data between different nodes and the business attributes carried by the data during the flow process; The converting the graphic description language file into a system architecture diagram further includes: Generate data flow labels between two nodes that have a connection relationship.

8. The method according to any one of claims 1 to 4, characterized in that: The system architecture diagram is included in the architecture document; The provider of the architecture document provides an update interface for the architecture document; the method further includes: Regularly obtain the latest submitted hash value of the target code from the remote repository of the software system, and compare the latest submitted hash value with the locally stored hash value; if the two are different, obtain the updated target code from the remote repository based on the latest submitted hash value; Inputting the updated target code into the large language model to obtain an updated graphic description language file output by the large language model; converting the updated graphic description language file into an updated system architecture diagram to obtain an updated version of the system architecture diagram; The update interface is called to replace the system architecture diagram in the architecture document with the updated version of the system architecture diagram.

9. A device for generating a system architecture diagram, characterized in that: The device comprises: A receiving module, used to receive the target code of the software system; the target code includes: components of the software system, and calling relationships and / or data flow relationships between components; A description language file determination module is used to determine a graph description language file corresponding to the target code using a large language model; the graph description language file is used to describe a graph containing nodes; wherein the nodes correspond to components in the target code; the graph description language file includes: a connection relationship between nodes; the connection relationship is used to characterize a calling relationship and / or a data flow relationship between components; A text-to-image conversion module, used to convert the graphic description language file into a system architecture diagram; An output module is used to output the system architecture diagram.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 8 when executing a program stored in a memory.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.