Cost-performance collaborative awareness micro-service identification method

Through the cost-performance collaborative perception microservice recognition method, the data center monolithic architecture is transformed into microservices, solving the problem of difficult to balance migration costs and division quality, and achieving high-performance and low-cost microservice recognition results.

CN120105056APending Publication Date: 2025-06-06BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510167556.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the process of microservice in the data center monolithic architecture, it is difficult to achieve a comprehensive balance between migration cost and division quality, and the relationship between system call characteristics and migration cost is not fully considered.

Method used

A microservice identification method for cost-performance collaborative perception is proposed. By analyzing the dynamic call behavior between methods in the data center log, a similarity matrix combining cost and performance characteristics is constructed, a hierarchical clustering algorithm is used to divide microservices, and the method-level division results are transformed into class-level results through the principle of majority voting, and the division scheme is optimized to achieve lower migration costs and higher division quality.

Benefits of technology

In the process of microserviceization of data center single-unit architecture, the balance between lower migration costs and higher division quality is achieved, providing high-performance and low-cost microservice identification results, which are suitable for practical scenarios such as architecture transformation and resource planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105056A_ABST
    Figure CN120105056A_ABST
Patent Text Reader

Abstract

The invention discloses a cost-performance collaborative awareness micro-service identification method. The overall process is divided into five steps of initialization, similarity matrix construction, clustering model construction, special method identification model construction and special method class mapping relation construction. According to the method, calling behavior characteristics and dependency relationships related to cost and performance are extracted for single application execution logs, and the micro-service identification method is constructed by adopting a hierarchical clustering algorithm based on the extracted characteristics and dependency relationships. Through the micro-service identification method provided by the invention, high-performance and low-cost micro-service division can be effectively realized, and the resource utilization efficiency and the service quality of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software engineering, and in particular relates to a method for micro-service-based single-body applications in a data center. Background Art

[0002] Data centers are the information infrastructure of the Internet and related industries. The evolution of their software architecture directly affects the service capabilities and efficiency of data centers. Early data centers were dominated by monolithic architectures, which were characterized by the integration of all functional modules into one system. This architecture is simple in design and easy to deploy, but as the scale of applications expands, its disadvantages gradually emerge, such as poor scalability, complex maintenance, and high risk of single point failure. In order to solve these problems, the microservice architecture came into being. The microservice architecture splits applications into more fine-grained services, and services communicate through lightweight protocols such as REST or gRPC. This architecture emphasizes loose coupling and autonomy, and each service can be developed, deployed, and expanded independently, greatly improving development efficiency and system elasticity. Since the monolithic architecture is usually a highly coupled whole and lacks clear boundaries between modules, splitting it into independent microservices requires in-depth analysis of the original system to clarify business logic and module dependencies. This splitting process is not only time-consuming and labor-intensive, but may also lead to functional duplication or failure of key functions. Therefore, the microserviceization of monolithic architecture is a complex process.

[0003] To promote the transformation of monolithic architecture to microservices, it is necessary to first collect data from monolithic applications to fully understand their operating characteristics. Then, by analyzing the data, the key features of the monolithic application are extracted as the basis for microservice identification. To achieve accurate microservice splitting, clustering algorithms or deep learning methods can be used to gradually complete the microservice identification work, thus laying a solid foundation for architecture optimization.

[0004] The existing work on microservices of monolithic data center architecture has the following defects: 1) No comprehensive evaluation of migration cost and partition quality. In data center monolithic applications developed in languages ​​such as Java, Python, and C++, classes are the basic units of programs, and methods are the core units of function execution and interaction. The two play different roles in microservice identification. Existing class-level microservice identification methods can significantly reduce migration costs, but it is often difficult to achieve the optimal partition quality; while method-level microservice identification can achieve higher partition quality, but ignores the high migration cost caused by high-quality partitioning. 2) Focus on analyzing the performance of call behavior. Existing methods mainly focus on the performance of system call behavior and analyze the performance characteristics of direct and indirect calls. However, these methods do not evaluate call behavior from the perspective of cost, lack in-depth consideration of the relationship between system call characteristics and migration cost, resulting in difficulty in achieving a comprehensive balance between performance and cost when identifying microservices.

[0005] The present invention aims to divide the data center monolithic architecture into microservices by comprehensively considering its migration cost and partition quality. By analyzing the system's dependency structure and resource usage characteristics, the partitioning scheme is optimized to achieve lower migration costs while ensuring higher partitioning quality. The obtained optimized partitioning results can be applied to practical scenarios such as architecture transformation and resource planning. Summary of the invention

[0006] Based on the above problems, the present invention proposes a microservice identification method for cost-performance collaborative perception. Aiming at monolithic applications, the present invention extracts the dynamic calling behavior between methods in the data center log. Based on the extracted calling behavior, a similarity matrix combining cost and performance characteristics is constructed, and then a hierarchical clustering algorithm is used to obtain the initial method-level microservice division. Then, the method-level division results are converted into class-level results according to the majority voting principle, and some special methods will be generated in this process. Finally, a new attribution relationship is constructed for special methods based on the similarity matrix. The model constructed by this method can obtain high-performance, low-cost microservice identification results.

[0007] The cost-performance collaboratively perceived microservice identification method proposed in this method mainly consists of five steps: initialization, similarity matrix construction, clustering model construction, special method identification model construction, and special method class mapping relationship construction. In this method, there are the following important parameters: the number of clusters clusterNum of the clustering algorithm, and the weight coefficient α of the similarity matrix. clusterNum is 2-20, and α is 0-1.

[0008] Before executing this method, the required log data is read in and converted into a processable form.

[0009] (1) Initialization

[0010] Use the method call information in the log to initialize the data, and let the full set of attributes contained in the log be T = {t 1 ,t 2 …t F}, from which the attribute subset with method-method calling relationship is selected as The call matrix C is constructed according to the calling relationship between methods in MD, and the class attribution matrix A is constructed according to the attribution relationship between methods and classes in MD. Its form is as follows:

[0011]

[0012] The elements in C and A are defined as shown in formula (1) and formula (2):

[0013]

[0014] (2) Similarity matrix construction

[0015] 2.1) Construct similarity matrix S based on matrix C 1 In matrix C, we will extract the common dependency relationship (CDR) and transitive dependency relationship (TDR) from the perspective of cost, and the direct call relationship (DCR) and indirect call relationship (ICR) from the perspective of performance.

[0016] Direct call relationship: In a single system, if method m i Directly calls method m j , then method m i and method m j There is a direct calling relationship between them. Here, method m is used i and method m j The call frequency between them is used to quantify the relationship between them, that is,

[0017]

[0018] Call(m i ,m j )=p,Call(m j ,m i )=q#(4)

[0019] Among them, Call(m i ,m j ) represents method m i Call method m j The number of times is p, Call(m j ,m i ) represents method m j Call method m i The number of times is q, n i Representative method m i The total number of calls that occurred, n j Representative method m j The total number of calls that occurred.

[0020] Indirect call relationship: In a single system, if method m i After k indirect calls, we reach method m j , and k is greater than 1, then method m i and method m jThere is an indirect call between them, and the calculation formula is as follows:

[0021]

[0022] Call(m o ,m o+1 )=p o #(7)

[0023] Among them, m o and m o+1 is m i After k indirect calls, it reaches m j The two methods involved in a direct call in the call path, k is the path depth of the indirect call, Call(m o ,m o+1 ) represents method m o Call method m o+1 The number of calls when o , n o For method m o The total number of calls that occurred, Call(m o ,m o+1 ) / n o represents the path weight of the oth indirect call, ICR(m i ,m j ) l Represents method m in the lth path i and method m j The calculation result when an indirect call occurs, u represents method m i To method m j There are u different indirect call paths, v represents method m j To method m i There are v different indirect call paths. For example, method A calls method B, and method B calls method C. Then there is an indirect call relationship between method A and method C. Substituting into the above definition, we can see that method A is m i , method C is m j , call depth k is 2, method m o and m o+1 There are two sets of values, method A calling method B and method B calling method C.

[0024] Co-dependency: In a monolithic system, if method m i Method m d A direct call occurs, method m j Also for method m d If a direct call occurs, method m i and method m j There is a common dependency relationship between them, and the calculation formula is as follows:

[0025]

[0026] X=[x 1 ,x 2 ,…,x l ],Y=[y 1 ,y 2 ,…,y l ]#(10)

[0027]

[0028] Call(m i ,m d )=p,Call(m j ,m d )=q#(12)

[0029] Among them, Call(m i ,m d ) represents method m i Directly call method m d The number of times is p, Call(m j ,m d ) represents method m j Directly call method m d The number of times is q, X and Y are a set of vectors of length l, representing method m i and method m j At the same time, there are direct calls to l different methods, H(m i ,m j ) is the calculation formula of cross entropy, which is used to measure the method m i and method m j For example, if method A calls method C, and method B calls method C, then there is a common dependency between method A and method B. Substituting into the above definition, we can see that method A is m i , method B is m j , method C is m d .

[0030] Transitive dependencies: In a monolithic system, if method m i After x indirect calls to reach method m r , method m j After y indirect calls to reach method m r , and x and y are both greater than 1, then method m i and method m j There is a transitive dependency relationship between them, and the calculation formula is as follows:

[0031]

[0032] X=[x 1 ,x 2 ,…,x l ],Y=[y 1 ,y 2 ,…,y l ]#(14)

[0033]

[0034] ICR(m i ,m r )=s,ICR(m j ,m r )=t#(16)

[0035] Among them, ICR (m i ,m r ) represents method m i Indirect call method m r The ICR value at this time is s (Formula (5)), ICR (m j ,m r ) represents method m j Indirect call method m r The ICR value when is t (Formula (5)), X and Y are a set of vectors of length l, representing method m i and method m j At the same time, there are indirect calls to l different methods, H(m i ,m j ) is the calculation formula of cross entropy (Formula (9)), which is used to measure the method m i and method m j For example, if method A calls method C, method C calls method E, method B calls method D, and method D calls method E, then there is a transitive dependency relationship between method A and method B. Substituting into the above definition, we can see that method A is m i , method B is m j , method E is m r .

[0036] Construct similarity matrix S 1 According to the calculation formulas of the above four relationships, the value range of DCR and ICR is between 0 and 1, while CDR and TDR use cross entropy, so their range is between 0 and positive infinity. Therefore, it is necessary to normalize these four relationships separately and then sum them to construct the similarity matrix S 1 .

[0037] S 1 (m i ,m j )=Norm(DCR(mi ,m j ))+Norm(ICR(m i ,m j ))+Norm(CDR(m i ,m j ))+Norm(TDR(m i ,m j ))#(17)

[0038] Where Norm is the maximum and minimum normalization algorithm, DCR(m i ,m j ) represents method m i and method m j The direct call relationship between them, ICR(m i ,m j ) represents method m i and method m j Indirect call relationship between CDR(m i ,m j ) represents method m i and method m j The common dependence between i ,m j ) represents method m i and method m j Transitive dependencies between .

[0039] 2.2) Construct similarity matrix S based on matrix A 2 The calculation formula is as follows:

[0040]

[0041] Where n represents the number of columns in matrix A, q k For method m i The value of the kth column of matrix A, p k For method m j The value of the kth column of matrix A, X(m i ,m j ) is method m i and method m j The Euclidean similarity.

[0042] 2.3) Merge S 1 and S 2 features, construct the similarity matrix S.

[0043] S=S 1 +α×S 2 #(20)

[0044]

[0045] in, Represents the similarity matrix S 1 and S 2 α is the weight coefficient of the similarity matrix. When constructing the similarity matrix S, the parameter α is selected mainly to balance the influence of different matrices and dynamically adapt to the characteristics of actual data.

[0046] (3) Clustering model construction

[0047] 3.1) Cluster the similarity matrix S using the agglomerative hierarchical clustering algorithm. The number of clusters in the hierarchical clustering is clusterNum. Initially, each method is located in an independent partition. At this time, the method set is C = {c 1 ,c 2 ,…c n},c i The representative method c is located in the i-th partition.

[0048] 3.2) Traverse all partitions and calculate the similarity between two partitions according to formula (22) in each round. i | and |n j | represents the number of methods in the i-th partition and the j-th partition, and the numerator is the sum of the similarity values ​​between all methods in partition i and partition j calculated according to the similarity matrix S. In each round, the two partitions with the highest similarity are selected for merging.

[0049]

[0050] 3.3) Repeat step 3.2) until n is less than clusterNum, and the method-level clustering is completed to obtain the clustering result.

[0051] (4) Construction of special method identification model

[0052] 4.1) Convert the method-level clustering results in (3) into class-level clustering results.

[0053] 4.2) Traverse all clusterNum partitions. If all methods in a class are still in the same partition after clustering, then the partition where the method is located is the partition of the class.

[0054] 4.3) Get all classes that have not been partitioned. Traverse all clusterNum partitions and get the number of methods of these classes in each partition. According to the majority voting principle, select the partition where most methods are located as the partition where the class is located.

[0055] 4.4) After 4.2) and 4.3), the clustering results at the method level have been transformed into the clustering results at the class level. However, two special methods will be generated in this process - isolated methods and free methods. An isolated method means that there is at least one class in the partition where the method is located and there is no class affiliation relationship between the method and any class in the partition. A free method means that there is no complete class in the partition where the method is located. Next, we will design a class mapping relationship model for these two methods.

[0056] (5) Construction of class mapping relationship for special methods

[0057] 5.1) For the free method, we first use the similarity matrix S 1 Calculate the cosine similarity between it and all methods in each candidate partition (Formula (23)). Then, sum the similarities of all methods in each partition, and reallocate the free method to the target partition with the highest similarity based on the maximization principle. After the allocation is completed, the free method needs to be attributed. If the free method is reallocated to the partition of its original class, its status will be converted to a normal method and its affiliation with the class will be reestablished. Otherwise, the free method will be marked as an isolated method.

[0058]

[0059] Among them, S 1 (i) represents the similarity matrix S 1 The i-th row element, S 1 (j) represents the similarity matrix S 1 The j-th row element, m i represents the ith method among the free methods, m j represents a method in each candidate partition, cosine(m i ,m j ) is used to calculate method m i and method m j The cosine similarity of .

[0060] 5.2) For the isolation method, we first use the similarity matrix S 1 Calculate the cosine similarity between it and the internal methods of each class in the partition (Formula (23)). Then, accumulate the similarities of the internal methods of each class, and classify the isolated methods into the class with the highest similarity to establish the belonging relationship. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 Cluster platform-dependent approach for cost-performance co-aware microservice identification.

[0062] Figure 2 This is a schematic diagram of the present invention.

[0063] Figure 3 It is a flow chart of the present invention.

[0064] Figure 4 Flowchart constructed for clustering model.

[0065] Figure 5 Flowchart constructed for class mapping of special methods. DETAILED DESCRIPTION

[0066] The present invention is described below in conjunction with the accompanying drawings and specific embodiments.

[0067] The cost-performance collaboratively perceived microservice identification method proposed in the present invention is built on multiple connected servers and implemented by writing corresponding functions. Figure 1 It is a deployment diagram of the platform built by this method. The platform consists of multiple computer servers (platform nodes), which are connected through a network to store data and execute tasks in a distributed manner. The platform nodes are divided into three categories: including a management node, a computing node, and an execution node. The platform built by the method of the present invention includes five core software modules: resource management module, task receiving module, task processing module, data receiving module, and data processing module. Among them, the resource management module is responsible for sending related tasks to the task receiving module, sending related log data to the data receiving module, and collecting the execution results of the management task processing module and the data processing module, and is only deployed on the management node; the task receiving module is responsible for receiving the request sent by the resource management module, and the module is deployed on the execution node; the task processing module is responsible for processing related requests, and writing the request execution path to the log, and finally returning the log record to the resource management module, which is deployed on the execution node; the data receiving module is responsible for pulling the required log data, which needs to be deployed on the computing node; the data processing module is responsible for running the corresponding algorithm and returning the result to the resource management module, which is deployed on the computing node. The above three types of software modules are all deployed and run when the platform is started.

[0068] Figure 2 The architecture diagram of the method of the present invention is as follows. The present invention takes the dynamic call log of a single application as input, and first constructs the dynamic call relationship matrix of the method and the class attribution relationship matrix of the method according to the call log. Based on the dynamic call relationship matrix of the method, the call behavior characteristics of cost and performance are extracted, and the similarity matrix S is constructed. 1 , extract the correlation features between methods based on the class attribution relationship matrix of the methods, and construct the similarity matrix S 2 . 1 and S 2Merge and construct the similarity matrix S. Use agglomerative hierarchical clustering algorithm to cluster the similarity matrix S and obtain the clustering results at the method level. Map the clustering results at the method level to the class level through the majority voting principle, obtain all special methods at the same time, and divide them according to their characteristics. Use cosine similarity to construct the optimal class relationship mapping for all special methods and obtain the final microservice division.

[0069] Combine the following Figure 3 The general process of the invention describes the specific implementation method of the present method. In the present implementation method, the basic parameters are set as follows: the weight coefficient of the similarity matrix α = 0.11. In order to verify the wide applicability of the present method, we conducted systematic experiments under multiple groups of different cluster numbers (the number of clusters of hierarchical clustering clusterNum = 5-13), and took the average of the experimental results to evaluate its overall performance.

[0070] The specific implementation method can be divided into the following steps:

[0071] 1. Initialization

[0072] The log used in the present invention contains the track of the dynamic call path in the system, which consists of the call time, calling method, called method, calling class, and called class. By analyzing all the call tracks in the system, the direct call relationship matrix C between methods and the attribution relationship matrix A between methods and classes are constructed.

[0073] 2. Similarity matrix construction

[0074] 2.1) According to the direct call relationship matrix between methods and the attribution relationship matrix between methods and classes, complete the similarity matrix S 1 and S 2 's construction.

[0075] 2.2) For each method in the system, perform S 1 and S 2 Data filling.

[0076] 2.2.1) Calculate the values ​​of direct call relationship, indirect call relationship, common dependency relationship, and transitive dependency relationship between two methods based on the direct call relationship matrix between methods, and fill them into the similarity matrix S 1 In; calculate the similarity between the two methods according to the attribution relationship matrix between the method and the class and fill it into the similarity matrix S 2 middle.

[0077] 2.2.2) Repeat step 2.2.1) until the similarity calculation of all methods is completed.

[0078] 2.3) The similarity matrix S1 And the similarity matrix S 2 Merge into the similarity matrix S, and the similarity matrix coefficient α is 0.11.

[0079] 3. Clustering model construction

[0080] 3.1) Construct a hierarchical clustering algorithm model.

[0081] 3.2) Perform hierarchical clustering on the above similarity matrix S, and the number of clusters clusterNum of hierarchical clustering are 5, 7, 9, 11, and 13 respectively.

[0082] 3.2.1) Initialize the partitions of all methods;

[0083] 3.2.2) Traverse all partitions, calculate the similarity between the methods within each two partitions according to the similarity matrix S, and record the two partitions with the highest similarity;

[0084] 3.2.3) Merge the methods in the two partitions with the highest similarity into one partition;

[0085] 3.2.4) Repeat 3.2.1) to 3.2.3). When the number of existing partitions is less than clusterNum, clustering stops.

[0086] 4. Special method identification model construction

[0087] 4.1) According to the partition results of the methods and the relationship between the methods and the classes, all the methods that still belong to the same class after clustering are determined, and the partitions where these methods are located are regarded as partitions of the class.

[0088] 4.2) According to the principle that "the partition where most methods are located is the partition where the class is located", assign corresponding partitions to the remaining classes.

[0089] 4.3) Identify isolated and free methods from all methods.

[0090] 5. Construction of class mapping relationship for special methods

[0091] 5.1) Constructing the class mapping relationship of free methods

[0092] 5.1.1) According to the similarity matrix S 1 Compute the cosine similarity between the free method and all common methods in the partition.

[0093] 5.1.2) Sum the cosine similarities of all methods within each partition.

[0094] 5.1.3) Select the partition with the highest similarity as the partition for the free method.

[0095] 5.1.4) If the partition of the free method is consistent with the partition of the class it originally belongs to, then the free method will be converted to a normal method and the free method will establish an affiliation relationship with the class. Otherwise, the free method will be converted to an isolated method.

[0096] 5.1.5) Repeat 5.1.1) to 5.1.4) until all free methods have class mapping relationships established or are converted into isolated methods.

[0097] 5.2) Construct a class mapping of isolated methods.

[0098] 5.2.1) According to the similarity matrix S 1 Compute the cosine similarity between an isolated method and all common methods in its partition.

[0099] 5.2.2) Based on the class, the cosine similarities of all methods in the partition are summed.

[0100] 5.2.3) Select the class with the highest similarity to establish the affiliation relationship.

[0101] 5.2.4) Repeat 5.2.1) to 5.2.3) until class mapping relationships are established for all isolated methods.

[0102] According to the cost-performance collaboratively perceived microservice identification method proposed in the present invention, the inventor conducted relevant performance tests. The test results show that the method of the present invention is applicable to monolithic applications with execution logs. The method can be used to accurately transform monolithic applications into microservices.

[0103] The performance test compares this method with the existing microservice identification method to reflect the advantages of the method proposed in this invention in microservice-based monolithic applications. The comparison method is as follows:

[0104] (1)Mono2micro: a practical and effective tool for decomposing monolithic java applications to microservices

[0105] Mono2micro collects runtime tracing information in a monolithic system by constructing and running user cases, and then constructs a similarity matrix by parsing the direct call relationships, indirect call relationships, direct call patterns, and indirect call patterns in the data, and then identifies microservices through hierarchical clustering.

[0106] (2)Microservice extraction using graph deep clustering based on dualview fusion

[0107] GDC-DVF uses runtime tracing data to build a class call graph (structural dependency view), generates a business function view through a random walk algorithm and uniform random sampling, uses a graph encoder based on a graph attention adaptive residual network to learn a fused feature embedding representation, and finally clusters the fused features to make microservice extraction recommendations.

[0108] (3)Analysis of a many-objective optimization approach for identifying microservices from legacy systems

[0109] toMicroservices uses the NSGA-III algorithm to identify microservices for legacy systems based on five key criteria (coupling, cohesion, functional modularity, network overhead, and reuse) and generates multiple possible microservice architecture solutions.

[0110] (4)Magnet:Method-based approach using graph neural network for microservices identification

[0111] MAGNET builds a feature-rich dependency graph containing system methods and their call relationships through static and semantic analysis. Then, it uses graph neural networks (GNNs) to perform unsupervised clustering on the dependency graph and automatically identify microservices.

[0112] The performance tests were run on a computer with an Intel(R) Core(TM) i5-14400 CPU and 16GB DDR4 RAM.

[0113] Four indicators are used to measure the functional independence of the divided microservices: (i) ICP; (ii) IFN; (iii) NED; (iv) CHD; ICP, IFN, and NED measure the independence of microservices at the class level, and CHD measures the independence of microservices at the method level. Two indicators, SM and SMQ, are used to measure the modularity of the divided microservices. SM measures the degree of modularity at the class level, and SMQ measures the degree of modularity at the method level. One indicator, cost, is used to measure the cost of new references in the system.

[0114] 1.ICP (Interpartition Call Percentage)

[0115] ICP is used to measure the percentage of calls between two microservices during runtime. The calculation formula is:

[0116]

[0117] Among them, c ij It is the number of calls that occur when a class in microservice i calls a class in microservice j, and n is the final number of microservices. The lower the ICP value, the more independent the microservice functions are after division.

[0118] 2.SM (Structural Modularity)

[0119] SM measures the quality of the divided microservices by calculating the difference between the cohesion (scoh) and coupling (scop) of the divided microservices. The calculation formula is:

[0120]

[0121] in, μ i Represents the number of all calls in the i-th microservice, m i is the number of all classes in the i-th microservice, σ ij Represents the number of calls between the i-th microservice and the j-th microservice, and n is the final number of microservices. The higher the SM value, the higher the degree of modularity of the divided microservices.

[0122] 3.IFN (Interface Number)

[0123] IFN is used to measure the number of interfaces published by microservices. The calculation formula is:

[0124]

[0125] Among them, ifn i is the number of interfaces in the i-th microservice, and n is the final number of microservices. The smaller the IFN value, the higher the independence of the divided microservices.

[0126] 4.NED (Non-Extreme Distribution)

[0127] NED is used to measure whether there is an extreme distribution of the divided microservices, that is, the number of classes in a microservice is too large. The calculation formula is:

[0128]

[0129] If the number of classes in the i-th microservice is between 5 and 20, then k i is the number of classes in the microservice. Otherwise, k i is 0. N is the number of all classes in the system, n is the final number of microservices, and the value of NED should be as small as possible.

[0130] 5.CHD (Cohesion at Domain level)

[0131] CHD is used to measure the number of interfaces published by microservices at the domain level. The calculation formula is:

[0132]

[0133] Among them, f term (opr k ) is to obtain the method opr k All keywords of f dom (opr k ,opr m ) is the method opr k and method opr m The difference measure of keyword coverage between i | is the number of methods in the i-th microservice, and n is the final number of microservices. The higher the CHD value, the higher the independence of the divided microservices.

[0134] 6.SMQ (Structural Modularity Quality)

[0135] SMQ, like SM, is used to measure the quality of the divided microservices. The calculation formula is the same as SM (Formula (30)), except that SM uses the number of calls between classes, while SMQ uses the number of calls between methods. The higher the SMQ value, the higher the degree of modularity of the divided microservices.

[0136] 7.cost

[0137] Cost is used to calculate the newly added migration cost when refactoring a method into a class. The calculation formula is as follows:

[0138] cost=C clusted +C no-cluster ×2#(31)

[0139] Among them, cost cluster Computation is the additional cost within the microservice. no-clusterIt is the additional cost between microservices. The smaller the cost value, the lower the cost of actual deployment of microservices.

[0140] The performance test selects the dynamic call log of Acmeair workload. The evaluation results at the class level (performance) are shown in Table 1. The evaluation results at the method level (cost) are shown in Table 2.

[0141] Table 1 Comparison results of microservice identification methods at the class level (performance)

[0142]

[0143] Table 2 Comparison results of microservice identification methods at the method level (cost)

[0144]

[0145] From the data in Table 1 and Table 2, it can be concluded that, compared with the comparison method, the method of the present invention has a higher performance improvement at the class level, and greatly reduces the cost at the method level while improving the performance. In summary, the method of the present invention improves the quality of microservice partitioning at the class level while reducing the actual code transformation cost.

[0146] Finally, it should be noted that the above examples are only used to illustrate the present invention and are not intended to limit the technology described in the present invention. All technical solutions and improvements that do not deviate from the spirit and scope of the invention should be included in the scope of the claims of the present invention.

Claims

1. A cost-performance collaboratively aware microservice identification method, characterized in that It consists of five steps: initialization, similarity matrix construction, clustering model construction, special method identification model construction, and special method class mapping relationship construction; it has the following parameters: the number of clusters clusterNum of the clustering algorithm, the weight coefficient α of the similarity matrix; clusterNum is 2-20, and α is 0-1; (1) Initialization Use the method call information in the log to initialize the data, and let the full set of attributes contained in the log be T = {t1, t2…t F }, from which the attribute subset with method-method calling relationship is selected as MD = {md1,md2…md S }, The call matrix C is constructed according to the calling relationship between methods in MD, and the class attribution matrix A is constructed according to the attribution relationship between methods and classes in MD. Its form is as follows: The elements in C and A are defined as shown in formula (1) and formula (2): (2) Similarity matrix construction 2.1) Construct a similarity matrix S1 based on the matrix C; extract the common dependency relationship (CDR) and the transitive dependency relationship (TDR) from the perspective of cost, and extract the direct call relationship (DCR) and the indirect call relationship (ICR) from the perspective of performance; Direct call relationship: In a single system, if method m i Directly calls method m j , then method m i and method m j There is a direct calling relationship between them. Here, method m is used i and method m j The call frequency between them is used to quantify the relationship between them, that is, Call(m i ,m j )=p,Call(m j ,m i )=q#(4) Among them, Call(m i ,m j ) represents method m i Call method m j The number of times is p, Call(m j ,m i ) represents method m j Call method m i The number of times is q, n i Representative method m i The total number of calls that occurred, n j Representative method m j The total number of calls that occurred; Indirect call relationship: In a single system, if method m i After k indirect calls, we reach method m j , and k is greater than 1, then method m i and method m j There is an indirect call between them, and the calculation formula is as follows: Call(m o ,m o+1 )=p o #(7) in, m o and m o+1 is m i After k indirect calls, it reaches m j The two methods involved in a direct call in the call path, k is the path depth of the indirect call, Call(m o ,m o+1 ) represents method m o Call method m o+1 The number of calls when o , n o For method m o The total number of calls that occurred, Call(m o ,m o+1 ) / n o represents the path weight of the oth indirect call, ICR(m i ,m j ) l Represents method m in the lth path i and method m j The calculation result when an indirect call occurs, u represents method m i To method m j There are u different indirect call paths, v represents method m j To method m i There are v different indirect call paths; For example, method A calls method B, and method B calls method C. Then there is an indirect calling relationship between method A and method C. Substituting this into the above definition, we can see that method A is m i , method C is m j , call depth k is 2, method m o and m o+1 There are two sets of values, method A calling method B, and method B calling method C; Co-dependency: In a monolithic system, if method m i Method m d A direct call occurs, method m j Also for method m d If a direct call occurs, method m i and method m j There is a common dependency relationship between them, and the calculation formula is as follows: H(m i ,m j )=-∑ l X(l)logY(l)#(9) X=[x1,x2,…,x l ],Y=[y1,y2,…,y l ]#(10) Call(m i ,m d )=p,Call(m j ,m d )=q#(12) Among them, Call(m i ,m d ) represents method m i Directly call method m d The number of times is p, Call(m j ,m d ) represents method m j Directly call method m d The number of times is q, X and Y are a set of vectors of length l, representing method m i and method m j At the same time, there are direct calls to l different methods, H(m i ,m j ) is the calculation formula of cross entropy, which is used to measure the method m i and method m j similarity; for example, method A calls method C, and method B calls method C, then there is a common dependency relationship between method A and method B; substituting into the above definition, we can see that method A is m i , method B is m j , method C is m d ; Transitive dependencies: In a monolithic system, if method m i After x indirect calls to reach method m r , method m j After y indirect calls to reach method m r , and x and y are both greater than 1, then method m i and method m j There is a transitive dependency relationship between them, and the calculation formula is as follows: X=[x1,x2,…,x l ],Y=[y1,y2,…,y l ]#(14) ICR(m i ,m r )=s,ICR(m j ,m r )=t#(16) Among them, ICR (m i ,m r ) represents method m i Indirect call method m r The ICR value at this time is s (Formula (5)), ICR (m j ,m r ) represents method m j Indirect call method m r The ICR value when is t (Formula (5)), X and Y are a set of vectors of length l, representing method m i and method m j At the same time, there are indirect calls to l different methods, H(m i ,m j ) is the calculation formula of cross entropy (Formula (9)), which is used to measure the method m i and method m j similarity; for example, method A calls method C, method C calls method E, method B calls method D, and method D calls method E, then there is a transitive dependency relationship between method A and method B; substituting into the above definition, we can see that method A is m i , method B is m j , method E is m r ; Constructing the similarity matrix S1: According to the calculation formulas of the above four relationships, the value range of DCR and ICR is between 0 and 1, while the value range of CDR and TDR is between 0 and positive infinity due to the use of cross entropy. Therefore, it is necessary to normalize these four relationships separately and then sum them to construct the similarity matrix S1; S1(m i ,m j )=Norm(DCR(m i ,m j ))+Norm(ICR(m i ,m j ))+ Norm(CDR(m i ,m j ))+Norm(TDR(m i ,m j ))#(17) Where Norm is the maximum and minimum normalization algorithm, DCT(m i ,m j ) represents method m i and method m j The direct call relationship between them, ICR (m i ,m j ) represents method m i and method m j Indirect call relationship between CDR(m i ,m j ) represents method m i and method m j The common dependence between i ,m j ) represents method m i and method m j Transitive dependencies between 2.2) Construct similarity matrix S2 based on matrix A; the calculation formula is as follows: Where n represents the number of columns in matrix A, q k For method m i The value in the kth column of matrix A, p k For method m j The value of the kth column of matrix A, X(m i ,m j ) is method m i and method m j The Euclidean similarity of 2.3) Combine the features of S1 and S2 to construct the similarity matrix S; S=S1+α×S2#(20) in, and Represents the mean of the similarity matrices S1 and S2; α is the weight coefficient of the similarity matrix; when constructing the similarity matrix S, the parameter α is selected mainly to balance the influence of different matrices and dynamically adapt to the characteristics of the actual data; (3) Clustering model construction 3.1) The similarity matrix S is clustered by agglomerative hierarchical clustering algorithm. The number of clusters in hierarchical clustering is clusterNum. When initialized, each method is located in an independent partition. At this time, the set of methods is C = {c1, c2, ... c n }, c i Representative method c is located in the i-th partition; 3.2) Traverse all partitions and calculate the similarity between two partitions in each round according to formula (22); where, |n on the denominator i | and |n j | represents the number of methods in the i-th partition and the j-th partition, and the numerator is the sum of the similarity values ​​between all methods in partition i and partition j calculated according to the similarity matrix S; in each round, the two partitions with the highest similarity are selected for merging; 3.3) Repeat step 3.2) until n is less than clusterNum, and the method-level clustering is completed to obtain the clustering result; (4) Construction of special method identification model 4.1) Convert the method-level clustering results in (3) into class-level clustering results; 4.2) Traverse all clusterNum partitions. If all methods in a class are still in the same partition after clustering, then the partition where the method is located is the partition of the class. 4.3) Get all classes that have not been partitioned; traverse all clusterNum partitions and get the number of methods of these classes in each partition; select the partition where most methods are located as the partition where the class is located according to the majority voting principle; 4.4) After 4.2) and 4.3), the clustering results at the method level have been transformed into the clustering results at the class level; however, two special methods will be generated in this process - isolated methods and free methods; isolated methods refer to methods where there is at least one class in the partition where the method is located and there is no class affiliation relationship between the method and any class in the partition; free methods refer to methods where there is no complete class in the partition where the method is located; next, we will design a class mapping relationship model for these two methods; (5) Construction of class mapping relationship for special methods 5.1) For the free method, first use the similarity matrix S1 to calculate the cosine similarity between it and all methods in each candidate partition (Formula (23)); Then, the similarities of all methods in each partition are summed, and based on the maximization principle, the free methods are reallocated to the target partition with the highest similarity; After the assignment is completed, the free method needs to be attributed. If the free method is reassigned to the partition of its original class, its status will be converted to a common method and its affiliation with the class will be reestablished. Otherwise, the free method will be marked as an isolated method. Among them, S1(i) represents the i-th row element of the similarity matrix S1, S1(j) represents the j-th row element of the similarity matrix S1, and m i represents the ith method among the free methods, m j represents a method in each candidate partition, cosine(m i ,m j ) is used to calculate method m i and method m j The cosine similarity of 5.2) For isolated methods, first use the similarity matrix S1, i.e., formula (23), to calculate the cosine similarity between the isolated method and the internal methods of each class in the partition. Then, add up the similarities of the internal methods of each class, and classify the isolated method into the class with the highest similarity to establish the affiliation relationship.