A microservice extraction method, system, medium, device and information processing terminal

By reverse engineering, the timing chart and class diagram are generated, and the microservice extraction is extracted in combination with the spectral clustering algorithm, the problems of high threshold and poor quality of microservice extraction in the existing technology are solved, and high-quality and automated microservice extraction is achieved.

CN115309634BActive Publication Date: 2025-07-01NORTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210835089.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-16
Publication Date
2025-07-01
Estimated Expiration
2042-07-16

AI Technical Summary

Technical Problem

The existing microservice extraction methods rely on test case design, with high threshold; source code-based methods cannot fully reflect business relationships, and the extracted microservice quality and availability are poor.

Method used

Time sequence diagrams and class diagrams are generated by reverse engineering, business functions are modeled based on class diagrams and timing diagrams, candidate microservices are extracted using spectral clustering algorithms, and their quality is evaluated and visualized.

Benefits of technology

It lowers the threshold for microservice extraction, improves the quality and availability of microservices, realizes the automation of microservice extraction, and reduces the time and labor costs of migration of single-unit architecture to microservice architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115309634B_ABST
    Figure CN115309634B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of microservice extraction, and discloses a microservice extraction method, system, medium, device and information processing terminal. The source code is hierarchically divided; the sequence diagram of each method in the control layer is obtained through reverse engineering, and the class diagrams of the entity layer, database access layer and other layers are obtained; the display business function modeling is carried out on the sequence diagram; the implicit business function modeling is carried out on the class diagram; the candidate microservices of the source code are extracted based on the business function model through spectral clustering; the quality of the candidate microservices is evaluated; the candidate microservices are visualized in the form of a graph structure to provide the architecture personnel with adjustment functions. The present invention uses the spectral clustering algorithm for microservice extraction, achieving the goal of high cohesion within microservices and low coupling between microservices; the extraction of microservices after the business function modeling of the source code greatly reduces the usage threshold of the method and reduces the time cost and labor cost of extracting from the monolithic architecture to microservices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microservice extraction, and in particular relates to a microservice extraction method, system, medium, device and information processing terminal. Background Art

[0002] Migrating a monolithic architecture system to a microservice architecture can better solve the problems faced by the monolithic architecture system. The task of microservice extraction is mainly to reasonably extract classes with the same business functions from the monolithic architecture system to form a group of candidate microservices. Compared with the traditional manual extraction and recognition methods, the existing microservice extraction methods achieve automatic or semi-automatic microservice recognition. However, the key input part that determines the quality of the extracted candidate microservices still depends on the generation of professionals, and the usage threshold of the method is relatively high, and the difficulty of the automatic extraction method is relatively high. Currently, the automatic extraction methods are mainly divided into two types: static and dynamic, with the following characteristics: 1. The method of extracting microservices based on the program execution trace depends on the design of test cases, and the test cases must be generated by software designers who are familiar with the system. Legacy systems often have problems such as complex business, large turnover of developers, and incomplete documentation. Therefore, it is very difficult to design test cases to fully cover all business functions, and the incomplete design of test cases will lead to a low coverage rate of the proposed microservice classes. Therefore, the usage threshold of this method is relatively high. 2. The method of extracting microservices based on source code generally uses code similarity as the basis for microservice extraction, which cannot fully reflect the business relationship, and the quality of the extracted microservices is poor and the usability is poor. The quality of microservices will largely determine the scalability of microservices. Therefore, there are relatively large hidden dangers in the use of this method. The existence of the above problems has led to the difficulty of automatic microservice extraction.

[0003] Through the above analysis, the problems and defects existing in the prior art are as follows: The method of extracting microservices based on the program execution trace depends on the design of test cases, and the test cases must be generated by software designers who are familiar with the system. Legacy systems often have problems such as complex business, large turnover of developers, and incomplete documentation. Therefore, it is very difficult to design test cases to fully cover all business functions, and the incomplete design of test cases will lead to a low coverage rate of the proposed microservice classes. Therefore, the usage threshold of this method is relatively high. The method of extracting microservices based on source code generally uses code similarity as the basis for microservice extraction, which cannot fully reflect the business relationship, and the quality of the extracted microservices is poor and the usability is poor. The quality of microservices will largely determine the scalability of microservices. Therefore, the applicability of this method is not high. Summary of the Invention

[0004] In view of the problems existing in the prior art, the present invention provides a microservice extraction method, system, medium, device and information processing terminal, and particularly relates to a microservice extraction method, system, medium, device and information processing terminal based on the source code of a monolithic system.

[0005] The present invention is implemented as follows. A microservice extraction method, the microservice extraction method comprising: generating a class diagram and a sequence diagram based on the source code; extracting business functions based on the class diagram and the sequence diagram; using a fast spectral clustering algorithm to perform microservice extraction on the extracted business functions, and finally performing quality evaluation on the generated candidate microservices and visually displaying them.

[0006] Performing hierarchical partitioning on the source code; obtaining the sequence diagram of each method in the control layer through reverse engineering, and obtaining the class diagrams of the entity layer, the database access layer and other layers; performing display business function modeling on the sequence diagram; performing implicit business function modeling on the class diagram; extracting candidate microservices of the source code based on the business function model through spectral clustering; evaluating the quality of the candidate microservices; and finally visualizing the candidate microservices in the form of a graph structure to provide an adjustment function for the architecture personnel.

[0007] Furthermore, the microservice extraction method comprises the following steps:

[0008] Step 1, performing hierarchical partitioning on the source code, obtaining the sequence diagram of the control layer through reverse engineering, and obtaining the class diagrams of the entity layer, the database access layer and other layers; because non-standard source code writing will affect the extraction of microservices.

[0009] Step 2, performing display business function modeling on the sequence diagram and performing implicit business function modeling on the class diagram; because the sequence diagram can reflect the dynamic relationship between classes, and analyzing the class diagram mainly considers that the domain-driven model is a relatively popular idea for current microservice splitting. Through class diagram analysis, domain relationships can be mined, thereby supplementing some business relationships that the sequence diagram cannot reflect.

[0010] Step 3, extracting candidate microservices of the source code based on the business function model through spectral clustering, because spectral clustering regards the identification of microservices as graph clustering and extracts microservices with high cohesion and low coupling.

[0011] Step 4, evaluating the quality of the candidate microservices, visualizing the candidate microservices in the form of a graph structure and providing an adjustment function; because using microservices with poor quality to improve legacy systems has poor scalability and cannot fully reflect the advantages of the microservice architecture. Evaluating the quality of microservices can avoid this situation.

[0012] Furthermore, in the step 1, performing hierarchical partitioning on the source code of the monolithic legacy system, and identifying the control layer, the business layer, the database access layer, the entity layer and other layers in the source code.

[0013] Generate the sequence diagram and class diagram of the system through reverse engineering based on the source code of the monomer system program:

[0014] (1) Use the SequenceDiagram tool to generate a sequence diagram for each method in the control layer of the source code, where the control layer of the source code refers to the class that interacts with the user interface display;

[0015] (2) Use the PlantUML Parser tool to generate class diagrams for the database access layer, entity layer, and other layers for each class file of the source code.

[0016] Furthermore, in the second step, by performing display business function modeling on the sequence diagram, a mapping table of the call relationships between classes is obtained, specifically including:

[0017] (1) In multiple sequence diagram files, read a sequence diagram file for parsing. When the sequence diagram file is empty, end the display business function modeling process;

[0018] (2) Count the number of calls f i between two classes C j . Each time a call relationship between C ij and C i to C j appears, f ij is incremented by 1;

[0019] (3) Store the number of calls between two classes C i and C j in a Map structure, where the key is a string formed by concatenating the class names of the two classes with "_", and the value is the number of calls f ij ;

[0020] (4) After parsing all sequence diagram files, output the model Map.

[0021] Perform implicit business function modeling based on the class diagram to obtain a semantic similarity relationship matrix between classes:

[0022] (1) In multiple class diagram files, read two class diagrams for parsing. When the class diagram file is empty, end the implicit business function modeling process;

[0023] (2) Calculate the semantic similarity S i between two class diagrams C j and C ij; Create a dictionary of word bags based on the input text information. Match the words in the text with the keys in the word bag to obtain the corpus. Initialize the tf-idf transformation model to get the transformed corpus corpus_tfidf. Train the corpus_tfidf using the Lsi model and calculate the sparse matrix similarity. Convert the tokenized list for which similarity is to be found into the corpus doc_test_vec to obtain the text similarity.

[0024] Further, in the third step, after obtaining the business function model of the source code, use the spectral clustering algorithm for clustering to obtain candidate microservices.

[0025] (1) Construct a similarity matrix, the formula is as follows:

[0026]

[0027] Among them, W ij represents the similarity matrix. If there is a call relationship between two classes i and j, then w ij = map ij Conversely, w ij = S ij ; D ij is a diagonal matrix, and the values on the diagonal are the sum of the corresponding rows or columns in the W matrix;

[0028] (2) Construct the Laplacian matrix L and normalize L. The formula is as follows:

[0029] L = D - W;

[0030]

[0031] (3) Perform eigenvalue decomposition on the Laplacian matrix, use the Lanczos method to accelerate the decomposition process, obtain the first k smallest eigenvalues and the corresponding eigenvectors, and finally form a k-dimensional eigenmatrix F;

[0032] (4) Use Kmeans to cluster the k-dimensional eigenmatrix F; Use AFK-MC 2 to select the initial clustering center. Randomly select an initial center sample c1; Calculate the proposal distribution q(x) of all data sets, randomly draw a data point from q(x) and calculate the distance dx; Use Markov chain Monte Carlo sampling to generate a sequence of length m, and take the last k - 1 as the center points C = {C1, C2,... Ck};

[0033]

[0034]

[0035]

[0036] Use A-means to reduce the time for the K-means algorithm to assign data points to the cluster C; calculate the distance from each data point x i to all cluster centers, and select the centroid C closest to itself k , and calculate the equidistant index α of the point i , the formula is as follows:

[0037] α i = abs(||i - μ1|| 2 - ||i - μ2|| 2 );

[0038] Calculate the improvement threshold , the formula is as follows:

[0039]

[0040] When , x i will not move in the round iteration, and x i is fixedly assigned to the cluster C k , and will no longer participate in distance calculation and reallocation.

[0041] After spectral clustering, a set of clustering results is obtained, and each cluster is a microservice candidate.

[0042] Furthermore, the quality evaluation of the extracted candidate microservices in step four includes:

[0043] Using the method of modular evaluation, the higher the modular evaluation value, the lower the coupling degree between the microservices and the higher the aggregation degree within the microservices. The calculation formula is as follows:

[0044]

[0045] Among them, N is the number of microservices, N i is the number of classes within the microservice, W i is the sum of the weights between the internal edges of microservice i, and W i,j is the sum of the weights between microservice i and microservice j.

[0046] Visualize the extracted candidate microservices, and display the microservices to which each class file belongs and the call relationships between classes in the form of a graph structure based on Echarts, and manually adjust the microservices to which the classes belong by dragging the vertices.

[0047] Another object of the present invention is to provide a microservice extraction system applying the microservice extraction method described above. The microservice extraction system includes:

[0048] A hierarchical division module, which is used to divide the source code hierarchically, obtain the sequence diagram of the control layer through reverse engineering, and obtain the class diagrams of the entity layer, the database access layer, and other layers;

[0049] A business function modeling module, which is used to perform display business function modeling on the sequence diagram and perform implicit business function modeling on the class diagram;

[0050] A candidate microservice extraction module, which is used to extract candidate microservices of the source code based on the business function model through spectral clustering;

[0051] A quality evaluation module, which is used to evaluate the quality of the candidate microservices, visualize the candidate microservices in the form of a graph structure, and provide an adjustment function.

[0052] Another object of the present invention is to provide a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the microservice extraction method.

[0053] Another object of the present invention is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the microservice extraction method.

[0054] Another object of the present invention is to provide an information data processing terminal, which is used to implement the microservice extraction system.

[0055] Combined with the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solution to be protected by the present invention from the following aspects:

[0056] First, aiming at the technical problems existing in the above-mentioned prior art and the difficulty of solving this problem, closely combining the technical solution to be protected by the present invention and the results and data in the R & D process, etc., analyze in detail and deeply how the technical solution of the present invention solves the technical problems and the creative technical effects brought after solving the problems. The specific description is as follows:

[0057] The microservice extraction method provided by the present invention performs business function modeling based on source code, solving the problem of poor modularity in the method of directly extracting microservices based on source code similarity; uses the spectral clustering algorithm for microservice extraction, better achieving the goal of high cohesion within microservices and low coupling between microservices; greatly reducing the usage threshold of the method by performing microservice extraction after business function modeling on the source code. Architecture personnel can obtain candidate microservices without understanding the legacy system, realizing the automation of microservice extraction and reducing the time cost and labor cost of migrating from a monolithic architecture to microservices.

[0058] By proposing a method for business function modeling based on source code, the present invention solves the problem that the generation of the key input part of microservice extraction technology requires the participation of professionals and the usage threshold of the method is high. The invention enables architects to quickly obtain candidate microservices by automatically generating sequence diagrams and class diagrams through reverse engineering when performing microservice extraction on a legacy system of a monolithic architecture. There is no need to analyze the business functions of the legacy system, reducing the usage threshold of the microservice extraction method. The invention also visualizes the extracted microservices in the form of a graph, facilitating architects to intuitively observe the extraction results for adjustment.

[0059] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are specifically described as follows:

[0060] The microservice extraction method provided by the present invention is committed to using the source code of a monolithic legacy system to automatically obtain business function information for microservice extraction, reducing the usage threshold of the microservice extraction method. The invention helps architecture personnel quickly perform microservice extraction on a monolithic legacy system, enabling them to do so without understanding the business information of the monolithic legacy system, reducing the usage threshold of the method, facilitating technicians to obtain high-cohesion and low-coupling microservices more quickly, and reducing the cost and risk of migrating from a monolithic architecture to a microservice architecture.

[0061] Third, as auxiliary evidence of the creativity of the claims of the present invention, it is also reflected in the following important aspects:

[0062] (1) The technical solution of the present invention fills the domestic and foreign industry technical gaps:

[0063] The current mainstream microservice extraction technology has the following problems: 1. The microservice extraction method based on program execution trajectory relies on the design of test cases. Test cases must be generated by software designers who are familiar with the system. Legacy systems often have complex business, high developer turnover, and incomplete documentation. Therefore, it is difficult to design test cases to fully cover all business functions. The incomplete design of test cases will lead to a low coverage rate of the proposed microservice classes, so the threshold for using this method is high. 2. The microservice extraction method based on source code generally uses code similarity as the basis for microservice extraction, which cannot fully reflect business relationships. The extracted microservices have poor quality and poor availability. The quality of microservices will largely determine the scalability of microservices, so the use of this method has great hidden dangers. The existence of the above problems has led to the difficulty of automatic extraction of microservices.

[0064] This paper proposes a technology for extracting business functions based on source code, using reverse engineering to generate a sequence diagram from the source code, and then obtains the call relationship that reflects the business functions of the program. While improving the quality of microservices, it does not require the participation of experts and reduces the difficulty of automatic extraction of microservices.

[0065] (2) The technical solution of the present invention solves a technical problem that people have been eager to solve but have never been able to solve successfully:

[0066] Microservice extraction technology has been committed to studying a method that can automatically split a monolithic system into microservices. The obtained microservices have good functional modularity, are easy to expand, develop and maintain, and can solve the problem that monolithic systems are difficult to maintain and expand. However, the current microservice extraction method cannot solve the above problems at the same time.

[0067] The method for extracting business functions based on source code of the present invention solves the technical difficulty that the current automatic acquisition of program business functions requires analysis based on program running trajectories, is inseparable from expert-designed test cases, and cannot be fully automated; it overcomes the technical difficulty that the quality of microservices extracted by traditional source code-based methods is poor. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0069] Figure 1 is a flow chart of a microservice extraction method provided by an embodiment of the present invention;

[0070] Figure 2 It is the flowchart of display service function modeling provided by the embodiment of the present invention;

[0071] Figure 3 It is the flowchart of implicit service function modeling provided by the embodiment of the present invention;

[0072] Figure 4 It is the result diagram of candidate microservice visualization provided by the embodiment of the present invention. Specific Embodiments

[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0074] In view of the problems existing in the prior art, the present invention provides a microservice extraction method, system, medium, device and information processing terminal. The following describes the present invention in detail with reference to the accompanying drawings.

[0075] I. Explanation of the embodiments. In order to enable those skilled in the art to fully understand how the present invention is specifically implemented, this part is an explanatory embodiment that expands on the technical solutions of the claims.

[0076] As Figure 1 shown, the microservice extraction method provided by the embodiment of the present invention includes the following steps:

[0077] S101, perform hierarchical partitioning on the source code;

[0078] S102, obtain the sequence diagram and class diagram of the control layer through reverse engineering;

[0079] S103, perform display service function modeling on the sequence diagram;

[0080] S104, perform implicit service function modeling on the class diagram;

[0081] S105, extract candidate microservices of the source code based on the service function model through spectral clustering;

[0082] S106, evaluate the quality of the candidate microservices;

[0083] S107, visualize the candidate microservices in the form of a graph structure and provide an adjustment function.

[0084] As a preferred embodiment, the microservice extraction method provided by the embodiment of the present invention specifically includes the following steps:

[0085] S1: Perform hierarchical partitioning on the source code of the monolithic legacy system, as Figure 2As shown in the figure, identify the control layer, business layer, database access layer, entity layer, and other layers in the source code. Generally, for source code that follows development specifications, the same layer is located in the same package, and the class suffix names are the same. The control layer is generally located in the package where classes such as Controller, Action, or View that directly receive user requests and respond to users are located. The database access layer is generally located in the package where Dao or classes that directly interact with the database are located. The service layer is generally located in the package where classes such as Service that interact with both the control layer and the database layer are located. The classes in the entity layer are generally used to encapsulate data and perform data transmission. Other classes mainly involve configuration classes and utility classes of the project.

[0086] S2: Generate sequence diagrams and class diagrams through reverse engineering. Specifically, this step includes the following sub-steps:

[0087] S21: Use the SequenceDiagram tool to generate a sequence diagram for each method in the control layer of the source code, where the control layer of the source code refers to the classes that interact with the user interface display.

[0088] S22: Use the PlantUMLParser tool to generate a class diagram for each class file in the source code.

[0089] S3: Model the display business functions based on the sequence diagrams. Figure 2 The following is the flow chart of the display business function modeling provided by the embodiment of the present invention. The detailed process is as follows:

[0090] S31: First, in multiple sequence diagram files, read a sequence diagram file for parsing. When the sequence diagram file is empty, end the display business function modeling process.

[0091] S32: Count the number of calls f i between two classes C j . Each time a call relationship from C ij to C i appears, f j is incremented by 1. ij

[0092] S33: Store the number of calls between two classes C i and C j in a Map structure, where the key is the string formed by concatenating the class names of the two classes with "_", and the value is the number of calls f ij .

[0093] S34: After parsing all sequence diagram files, output the model Map.

[0094] S4: Model the implicit business functions based on the class diagrams. Figure 3The flowchart of implicit function modeling provided by the embodiments of the present invention is as follows:

[0095] S41: First, read and parse two class diagrams from multiple class diagram files. When the class diagram file is empty, end the implicit business function modeling process;

[0096] S42: Calculate the semantic similarity S i between two class diagrams C j and C ij . First, create a bag of words dictionary based on the input text information, then match the words in the text with the keys in the dictionary to obtain the corpus. Initialize the tf-idf transformation model to get the transformed corpus corpus_tfidf. Train the corpus_tfidf using the Lsi model, calculate the sparse matrix similarity, and convert the tokenized list for which similarity needs to be found into the corpus doc_test_vec to obtain the similarity of the text.

[0097] S5: Use spectral clustering for clustering to obtain candidate microservices. The detailed process is as follows:

[0098] S51: Construct a similarity matrix. The formula is as follows. W ij represents the similarity matrix. If there is a call relationship between two classes i and j, then w ij =map ij Otherwise, w ij =S ij ; D ij is a diagonal matrix, and the values on the diagonal are the sums of the corresponding rows or columns in the W matrix.

[0099]

[0100] S52: Construct a Laplacian matrix L and normalize L. The formula is as follows:

[0101] L = D - W

[0102]

[0103] S53: Perform eigenvalue decomposition on the Laplacian matrix and use the Lanczos method to accelerate the decomposition process to obtain the first k smallest eigenvalues and the corresponding eigenvectors, and finally form a k-dimensional eigenmatrix F.

[0104] S54: Use Kmeans to cluster the k-dimensional eigenmatrix F. First, use AFK-MC 2Select the initial cluster center, randomly select an initial center sample c1, then calculate the proposal distribution q(x) of all data sets (as shown in Formula 5), ​​randomly select a data point from q(x) and calculate the distance dx (as shown in Formula 6); finally, use Markov Chain Monte Carlo (Metropolis-Hastings) to sample a sequence of length m, and take the last k-1 as the center point C = {C1, C2, ..., Ck}.

[0105]

[0106]

[0107] d x =d(x,C i-1 ) 2

[0108] Use A-means to reduce the time of K-means algorithm to assign data points to clusters C. Calculate each data point x i The distance to all cluster centers, select the centroid C closest to itself k , calculate the isometry index α of the point i , the formula is as follows:

[0109] α i =abs(||i-μ1|| 2 -||i-μ2|| 2 )

[0110] Calculating the improvement threshold The formula is as follows:

[0111]

[0112] when When xi is assumed to not move in this round of iteration, x i Fixed assignment to cluster C k , no longer participate in distance calculation and redistribution.

[0113] After the spectral clustering in the above steps, a set of clustering results will be obtained, and each cluster is a microservice candidate.

[0114] S6: Evaluate the quality of the extracted candidate microservices. Use modular evaluation. The higher the modular evaluation value, the lower the coupling between the microservices and the higher the aggregation within the microservices. The formula is as follows:

[0115]

[0116] Where N is the number of microservices, N iis the number of classes within a microservice, W i is the sum of weights between internal edges within microservice i, W i,j is the sum of weights of the edge between microservice i and microservice j.

[0117] S7: Visualize the candidate microservices extracted by this method. Based on Echarts, display the microservices to which each class belongs in the form of a graph structure. As Figure 4 shown, where the vertices of the graph represent class files, the colors of the vertices represent the microservices to which they belong, vertices with the same color belong to the same microservice, and the edges represent the call relationships between classes.

[0118] The microservice extraction system provided by the embodiments of the present invention includes:

[0119] A hierarchical division module, configured to perform hierarchical division on the source code, obtain the sequence diagram of the control layer through reverse engineering, and obtain the class diagrams of the entity layer, the database access layer, and other layers;

[0120] A business function modeling module, configured to perform display business function modeling on the sequence diagram and perform implicit business function modeling on the class diagram;

[0121] A candidate microservice extraction module, configured to extract candidate microservices of the source code based on the business function model through spectral clustering;

[0122] A quality evaluation module, configured to evaluate the quality of the candidate microservices, visualize the candidate microservices in the form of a graph structure, and provide an adjustment function.

[0123] II. Application embodiments. In order to prove the creativity and technical value of the technical solution of the present invention, this part is an application embodiment of the technical solution of the claims on a specific product or related technology.

[0124] The following describes the technical solution in detail in combination with the microservice extraction embodiment of the application program SpringBlog.

[0125] First step, pull the source code of the SpringBlog application program on github and analyze the hierarchical relationship of the source code. Identify the control layer, business layer, database access layer, entity layer, and other layers in the source code. The classes in the control layer end with the suffix Controller and are the classes that directly receive user requests and respond to users. The database access layer is generally located in repositories, in the package that interacts with the database. The service layer is located in packages such as Service, which interacts with both the control layer and the database layer. The classes in the entity layer are generally used to encapsulate data and perform data transmission, mainly located in the models and forms packages. Other classes mainly involve the configuration classes and tool classes of the project.

[0126] Step 2: Use tools to obtain the sequence diagram of the control layer of SpringBlog and the class diagrams of other layers.

[0127] Step 3: Perform display business function modeling based on the sequence diagram. In multiple sequence diagram files, read and parse one sequence diagram file. When the sequence diagram file is empty, end the display business function modeling process; output the model Map. Count the number of calls f i between class C j and C ij . Each time there is a call relationship from C i to C j , f ij is incremented by 1; store the number of calls between two classes C i and C j in the Map structure, where the key is the string formed by concatenating the class names of the two classes with "_", and the value is the number of calls f ij ;

[0128] Step 4: Perform implicit business function modeling based on the class diagram. Calculate the semantic similarity S ij between two class diagrams Ci and Cj. First, create a bag of words dictionary based on the input text information, then match the words in the text with the keys in the dictionary to obtain the corpus. Initialize the tf-idf transformation model to get the transformed corpus corpus_tfidf. Train the corpus_tfidf using the Lsi model, calculate the sparse matrix similarity, and perform format conversion to make the list of segmented words to be searched for similarity into the corpus doc_test_vec to obtain the similarity of the text.

[0129] Step 5: Use spectral clustering for clustering to obtain candidate microservices.

[0130] Step 6: Use the method of modular evaluation. The higher the modular evaluation value, the lower the coupling degree between the microservices and the higher the aggregation degree within the microservices. The formula is as follows:

[0131]

[0132] where N is the number of microservices, N i is the number of classes within the microservice, Wi i is the sum of the weights of the edges within microservice i, and Wij i,j is the sum of the weights of the edges between microservice i and microservice j.

[0133] Step 7: Visualize the candidate microservices extracted by this method.

[0134] III. Evidence of the effects related to the embodiments. Some positive effects have been achieved during the R & D or use of the embodiments of the present invention, and there are indeed great advantages compared with the prior art. The following content will be described in combination with the data, charts, etc. in the test process.

[0135] This paper proposes that compared with the traditional microservice extraction method based on source code, the module quality has indeed been greatly improved. The following table shows the quality evaluation of microservice extraction from the SpringBlog application by the two methods.

[0136] Among them, MEMSC is the MQ value of the microservices extracted by the traditional method, and OUR is the MQ value of the microservices obtained by the method proposed in this paper. K is the splitting granularity of the microservices. We split the SpringBlog application into 6 microservices, 7 microservices, 8 microservices, 9 microservices, and 10 microservices respectively, and statistically analyzed the module quality MQ of the microservices. The value of MQ represents the quality of the microservices. The higher this value is, the better the modularity of the microservices is, and the more the microservices meet the standard of high cohesion and low coupling.

[0137] It can be seen through observation that the microservice extraction method proposed in this aspect has higher microservice quality than the traditional microservice extraction method under five different splitting granularities. Especially when the splitting granularity k = 9, the quality of the microservices has been greatly improved.

[0138]

[0139] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code. For example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.

[0140] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A microservice extraction method, characterized in that, The microservice extraction method includes the following steps: Step 1: Divide the source code into layers, obtain the sequence diagram of the control layer through reverse engineering, and obtain the class diagrams of the entity layer, database access layer, and other layers; Step 2: Perform explicit business function modeling on the sequence diagram and implicit business function modeling on the class diagram; Step 3: Extract candidate microservices from the source code based on the business function model through spectral clustering; Step 4: Evaluate the quality of the candidate microservices, visualize the candidate microservices in the form of a graph structure, and provide an adjustment function; In the second step, through explicit business function modeling of the sequence diagram, a call relationship mapping table between classes is obtained, which specifically includes: (1) In multiple sequence diagram files, read and parse one sequence diagram file. When the sequence diagram file is empty, end the explicit business function modeling process; (2) Count the number of calls fij between two classes Ci and Cj. Each time a call relationship from Ci to Cj appears, fij is incremented by 1; (3) Store the number of calls between two classes Ci and Cj in a Map structure, where the key is a string formed by concatenating the class names of the two classes using "_", and the value is the number of calls fij; (4) After parsing all sequence diagram files, output the model Map; According to the class diagram, perform implicit business function modeling to obtain a semantic similarity relationship matrix between classes: (1) In multiple class diagram files, read and parse two class diagrams. When the class diagram file is empty, end the implicit business function modeling process; (2) Calculate the semantic similarity Sij between two class diagrams Ci and Cj; make a bag of words dictionary according to the input text information, match the words in the text with the keys in the bag of words to obtain a corpus; initialize the tf-idf transformation model to obtain the transformed corpus corpus_tfidf, train the corpus_tfidf corpus using the Lsi model, and calculate the sparse matrix similarity; perform format conversion to make the word segmentation list to be searched for similarity into a corpus doc_test_vec to obtain the similarity of the text; In the third step, after obtaining the business function model of the source code, use the spectral clustering algorithm for clustering to obtain candidate microservices; (1) Construct a similarity matrix, and the formula is as follows: Among them, Wij represents the similarity matrix. If there is a call relationship between two classes i and j, then wij = mapij, otherwise wij = Sij; Dij is a diagonal matrix, and the values on the diagonal are the sum of the corresponding rows or columns in the W matrix; (2) Construct a Laplacian matrix L and normalize L. The formula is as follows: L = D - W; (3) Perform eigenvalue decomposition on the Laplacian matrix, use the Lanczos method to accelerate the decomposition process, obtain the first k smallest eigenvalues and the corresponding eigenvectors, and finally form a k-dimensional eigenmatrix F; (4) Cluster the k-dimensional feature matrix F using Kmeans; use AFK-MC2 to select the initial cluster centers, and randomly select an initial center sample c1; calculate the proposal distribution q(x) of all data sets, randomly draw a data point from q(x) and calculate the distance dx; sample a sequence of length m using Markov chain Monte Carlo sampling, and take the last k-1 as the center points C = {C1, C2,..., Ck}; d x = d(x, C i-1 ) 2 ; Reduce the time taken by the K-means algorithm to assign data points to cluster C using A-means; calculate the distance of each data point x i to all cluster centers and select the centroid C that is closest to itself k , calculate the equidistant index α of the point i , the formula is as follows: α i = abs(||i - μ1|| 2 - ||i - μ2|| 2 ); Calculate the improvement threshold The formula is as follows: When , x i will not move during the round iteration, and x i is fixedly assigned to cluster C k , and will no longer participate in distance calculation and reallocation; After spectral clustering, a set of clustering results is obtained, and each cluster is a microservice candidate.

2. The microservice extraction method according to claim 1, wherein In the first step, hierarchically partition the source code of the monolithic legacy system, and identify the control layer, business layer, database access layer, entity layer, and other layers in the source code; Generate the sequence diagram and class diagram of the system through reverse engineering based on the source code of the monolithic system: (1) Use the SequenceDiagram tool to generate a sequence diagram for each method in the control layer of the source code, where the control layer of the source code refers to the classes responsible for interacting with the user interface display; (2) Use the PlantUML Parser tool to generate class diagrams of the database access layer, entity layer, and other layers for each class file of the source code.

3. The microservice extraction method according to claim 1, characterized in that, The quality evaluation of the extracted candidate microservices in step 4 includes: Use the modular evaluation method. The higher the modular evaluation value, the lower the coupling degree between the microservices and the higher the aggregation degree within the microservices. The calculation formula is as follows: Among them, N is the number of microservices, N i is the number of classes within a microservice, W i is the sum of weights between internal edges within microservice i, W i,j is the sum of weights of the edge between microservice i and microservice j; Visualize the extracted candidate microservices, display the microservices to which each class file belongs and the call relationships between classes in the form of a graph structure based on Echarts, and manually adjust the microservices to which the classes belong by dragging the vertices.

4. A microservice extraction system applying the microservice extraction method according to any one of claims 1 to 3, characterized in that The microservice extraction system includes: A hierarchical partitioning module for hierarchically partitioning the source code, obtaining the sequence diagram of the control layer through reverse engineering, and obtaining the class diagrams of the entity layer, database access layer, and other layers; A business function modeling module for performing display business function modeling on the sequence diagram and implicit business function modeling on the class diagram; A candidate microservice extraction module for extracting candidate microservices of the source code based on the business function model through spectral clustering; A quality evaluation module for evaluating the quality of the candidate microservices, visualizing the candidate microservices in the form of a graph structure, and providing an adjustment function.

5. A computer device, characterized in that, The computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the microservice extraction method according to any one of claims 1 to 3.

6. A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the microservice extraction method according to any one of claims 1 to 3.

7. An information data processing terminal, characterized in that The information data processing terminal is used to implement the microservice extraction system according to claim 4.