A software component identification method and system based on spectral clustering
Through a spectral clustering-based method, the class call similarity matrix of the software system is constructed and processed, and combined with the Discretize clustering algorithm and component quality function, the components in the software system are automatically identified, solving the problem of inaccurate identification and user input parameters in the existing technology, and efficient and automated component recognition is achieved.
Patent Information
- Application Number
- CN202310976538.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-04
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-08-04
AI Technical Summary
The prior art is difficult to accurately identify components in software systems, especially when processing high-dimensional and sparse data, and often users need to input parameters to obtain optimal solutions.
Using a spectral clustering-based method, the similarity matrix and Laplace matrix are constructed, the eigenvalue and eigenvector are calculated, and the number of clusters is automatically determined and the components are identified by combining the Discretize clustering algorithm and component quality function.
It realizes accurate identification of components in the software system without user input parameters, improves the recognition quality and automation of algorithms, and reduces the operation time.
Smart Images

Figure CN116933117B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of process mining, and relates to a software component recognition method and system based on spectral clustering. Background Art
[0002] Software component recognition refers to identifying independent, reusable, and functionally complete modules or components in a software system to support tasks such as software system maintenance, upgrade, reconstruction, and reuse. Component-based development can split the system into independent modules, enabling developers to focus more on module development and maintenance, thereby improving software development efficiency. By identifying and classifying components, the structure and function of the software system can be understood more clearly, making it easier to maintain and upgrade the software system. Standardizing and modularizing components allows different software systems to share the same components, thus improving the reusability of software systems.
[0003] In the field of software processes, component recognition can help us understand and improve various activities and tasks in the software development process, such as requirements analysis, design, coding, testing, and deployment. By performing data analysis and mining on these activities and tasks, we can identify potential problems and bottlenecks, and then take corresponding measures for optimization and improvement.
[0004] The analysis and understanding of the software system architecture can be used as a guide to help understand the operation and evolution of the software, which is beneficial for future software maintenance. A software system usually contains multiple interacting logical components. Components are the units of functional division in the software system. They are meaningful, almost independent, and replaceable. If the behavior inside components and the interactions between different components can be explored, a better understanding of the software system architecture can be achieved.
[0005] Spectral clustering is an unsupervised clustering algorithm based on graph theory. It realizes data clustering by representing data samples as nodes on a graph and performing clustering on the graph. Specifically, spectral clustering regards data samples as nodes on a graph and represents their similarity as the edge weights between nodes. Then, it calculates the Laplacian matrix of this graph and obtains the clustering result through eigenvalue decomposition of the Laplacian matrix. Spectral clustering can not only effectively handle data clusters with non-convex and complex shapes, but also has good scalability when dealing with large-scale data sets. In addition, spectral clustering does not require specifying the number of clusters in advance, but determines the number of clusters through the eigenvalues of the Laplacian matrix, so it is also very applicable when dealing with data sets with an uncertain number of clusters. Moreover, since spectral clustering is a graph theory-based method, it can naturally handle complex situations such as weighted graphs, data sets with noise, and multi-view data sets. Summary of the Invention
[0006] To solve the above problems, the present invention provides a software component recognition method and system based on spectral clustering.
[0007] In a first aspect, a software component recognition method based on spectral clustering includes the following steps:
[0008] S1. Obtain the software running event log SE, and obtain the set of all classes U involved in the software running event log SE Cl ;
[0009] S2. Construct a class call similarity matrix according to the software running event log SE and its set of classes U Cl ;
[0010] S3. Construct a degree matrix according to the class call similarity matrix, and calculate the Laplacian matrix based on the degree matrix;
[0011] S4. Calculate all the eigenvalues of the Laplacian matrix and sort them in ascending order, and construct an eigenvector space through the eigenvectors of the first k eigenvalues;
[0012] S5. Use the Discretize clustering algorithm to cluster the eigenvector space, and obtain the clustering result of the component with the highest quality through the component quality function as the component recognition result;
[0013] S6. Add component attribute information to the software running event log SE through the component recognition result.
[0014] Further, step S1 obtains the software running event log SE = {c 0 , c 1 , …, c M} containing M groups of software running cases, and obtains the set of all classes U involved in the software running event log SE Cl = {cl 1 , cl 2 , …, cl N}; where c m , m = 1, 2, …, M represents the m-th group of software running cases in the software running event log SE, and cl i , i = 1, 2, …, N represents the i-th class.
[0015] Further, the class call similarity matrix is obtained by conversion according to the software running event log SE and its set of classes U Cl , including:
[0016] S21. Treat each class as a node, and connect edges according to the interaction relationship between classes;
[0017] S22. Calculate the total number of interactions between each pair of classes, expressed as:
[0018]
[0019]
[0020] where W SE (cl i ,cl j ) represents the total number of interactions between class cl i and class cl j , W cm (cl 1 ,cl 2 ) represents the interaction result between class cl i and class cl j in the m-th group of software running cases c m . represents that in the m-th group of software running cases c m class cl i calls class cl j , represents that in the m-th group of software running cases c m class cl j calls class cl i ;
[0021] S23. Consider the total number of interactions between each pair of classes as the edge weight of its corresponding edge, and construct an N×N class call similarity matrix W.
[0022] Furthermore, the specific process of step S3 includes:
[0023] S31. Calculate the sum of the elements in each row of the class call similarity matrix, and use d i , i = 1, 2,..., N to represent the sum of the elements in the i-th row of the class call similarity matrix;
[0024] S32. Construct an N×N matrix D, and use D ij to represent the element in the i-th row and j-th column of matrix D; let D ii = d i and the remaining elements be 0, to obtain the degree matrix corresponding to the class call similarity matrix;
[0025] S33. Calculate the Laplacian matrix based on the degree matrix, and the calculation expression is:
[0026] Lap = De -1 (De - W)
[0027] where Lap represents the Laplacian matrix, De represents the degree matrix, and W represents the class call similarity matrix.
[0028] Further, in step S5, the Discretize clustering algorithm is used to cluster the feature vector space to obtain multiple clusters, and the quality of each cluster is evaluated by a component quality function, and the cluster with the best quality is selected as the component recognition result; the evaluation formula of the component quality function is:
[0029] ComQuality = (MQ + RIC) / 2
[0030] where MQ is the modular quality and RIC is the proportion of intermediate class components.
[0031] Further, the specific process of step S6 includes:
[0032] According to the clustering result, component attribute information is added to each group of software operation cases in the software event log.
[0033] In a second aspect, based on the method proposed in the first aspect, the present invention also provides a software component recognition system based on spectral clustering, including:
[0034] A data acquisition module, configured to acquire a software operation event log SE = {c 0 , c 1 , …, c M} containing M groups of software operation cases, and acquire all classes involved in the software operation event log SE to form a class set U Cl = {cl 1 , cl 2 , …, cl N}; where c m , m = 1, 2, …, M represents the mth group of software operation cases in the software operation event log SE, and cl i , i = 1, 2, …, N represents the ith class;
[0035] A class call similarity matrix construction module, configured to construct a class call similarity matrix according to the software operation event log SE and its class set U Cl ;
[0036] A feature vector space construction module, configured to construct a degree matrix according to the class call similarity matrix, and calculate a Laplacian matrix based on the degree matrix; calculate all eigenvalues of the Laplacian matrix and arrange them in ascending order, and construct a feature vector space through the eigenvectors of the first k eigenvalues;
[0037] A clustering output module, configured to cluster the feature vector space by using the Discretize clustering algorithm, and obtain the clustering result with the component of the highest quality as the component recognition result through the component quality function; after adding component attribute information to the software operation event log SE through this component recognition result, output it.
[0038] Advantages of the present invention:
[0039] The present invention proposes a component identification algorithm based on spectral clustering (A Software Execution datacomponent identification method based on Spectral Clustering, abbreviated as SESC). Compared with the existing algorithms for identifying components through software operation data, the algorithm proposed by the present invention can take into account the number of calls between classes to achieve a more accurate purpose of identifying components, and the SESC algorithm is the first application of spectral clustering in this field.
[0040] Compared with the existing component identification algorithms, the SESC algorithm obtains the optimal solution of spectral clustering by setting a component quality function and traversing all possible numbers of clusters, that is, the situation where the component quality function reaches the maximum value. This algorithm can automatically obtain the optimal solution without the need for users to input parameters.
[0041] To ensure the accuracy of the proposed algorithm, the quality of the identified components is evaluated according to quality criteria such as cohesion and coupling.
[0042] A large number of experiments have been carried out on real software operation data sets. Compared with other existing component identification algorithms, the algorithm proposed by the present invention has significant advantages in terms of the quality of the identified components and the running time of the algorithm. Brief Description of the Drawings
[0043] Figure 1 is a framework diagram of a component identification algorithm based on spectral clustering of the present invention;
[0044] Figure 2 is a comparison of the number of components identified by different algorithms in the embodiments of the present invention;
[0045] Figure 3 is the value of the ratio of singleclass components (abbreviated as RSC) and the ratio of intermediate components (abbreviated as RIC) identified by different algorithms in the embodiments of the present invention;
[0046] Figure 4 is the value of the component quality function ComQuality identified by different algorithms in the embodiments of the present invention;
[0047] Figure 5 is the time consumed by different algorithms in the embodiments of the present invention. Detailed Embodiments
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] The present invention provides a software component recognition method based on spectral clustering, which applies spectral clustering to the field of software recognition. The main reasons for doing so are as follows:
[0050] 1. Spectral clustering can effectively reduce the dimension of high-dimensional data, reduce it to a low dimension and then use a classical clustering algorithm for clustering, so it is more suitable for processing high-dimensional data;
[0051] 2. Spectral clustering converts data into a similarity matrix, and then clusters after constructing a feature vector space for the similarity matrix, which is very suitable for processing sparse data, while traditional clustering algorithms are difficult to handle sparse data well.
[0052] In one embodiment, the software component recognition method based on spectral clustering, as Figure 1 shown, includes the following steps:
[0053] S1. Obtain the software running event log SE, and obtain the set of all classes U involved in the software running event log SE Cl .
[0054] Specifically, step S1 obtains the software running event log SE = {c 0 , c 1 , …, c M} containing M groups of software running cases, and obtains the set of all classes U involved in the software running event log SE Cl = {cl 1 , cl 2 , …, cl N}; where c m , m = 1, 2, …, M represents the mth group of software running cases in the software running event log SE, and cl i , i = 1, 2, …, N represents the ith class.
[0055] S2. Construct a class call similarity matrix according to the software running event log SE and its class set U Cl .
[0056] Specifically, converting according to the software running event log SE and its class set U Cl to obtain a class call similarity matrix includes:
[0057] S21. Treat each class as a node and connect edges according to the interaction relationships between classes. That is, if there is an interaction between two classes, there is an edge connecting them; if there is no interaction between two classes, there is no edge connecting them, which also means there is no edge weight between these two classes.
[0058] S22. Calculate the total number of interactions between every two classes, which is expressed as:
[0059]
[0060]
[0061] Among them, W SE (cl i ,cl j ) represents the total number of interactions between class cl i and class cl j . represents the interaction result between class cl i and class cl j in the m-th group of software running cases c m . represents that in the m-th group of software running cases c m class cl i invokes class cl j . i represents that in the m-th group of software running cases c m class cl j invokes class cl i .
[0062] S23. Treat the total number of interactions between every two classes as the edge weight of its corresponding edge, and construct an N×N class call similarity matrix W.
[0063] Specifically, using the above method, a class call similarity matrix W as shown below can be obtained:
[0064]
[0065] S3. Construct a degree matrix based on the class call similarity matrix, and calculate the Laplacian matrix based on the degree matrix.
[0066] Specifically, the specific process of step S3 includes:
[0067] S31. Calculate the sum of the elements in each row of the class call similarity matrix, and use d i , i = 1, 2, …, N to represent the sum of the elements in the i-th row of the class call similarity matrix;
[0068] S32. Construct an N×N matrix D, and use Dij Denote the element in the \(i\)-th row and \(j\)-th column of matrix \(D\); let \(D\) ii = \(d\) i and the rest of the elements are 0, obtaining the degree matrix corresponding to the class call similarity matrix;
[0069] S33. Calculate the Laplacian matrix according to the degree matrix, and the calculation expression is:
[0070] Lap = \(D_e\) -1 (\(D_e - W\))
[0071] where Lap represents the Laplacian matrix, \(D_e\) represents the degree matrix, and W represents the class call similarity matrix.
[0072] S4. Calculate all the eigenvalues of the Laplacian matrix and sort them in ascending order, and construct the eigenvector space through the eigenvectors of the first \(k\) eigenvalues.
[0073] Specifically, the specific process of step S4 includes:
[0074] S41. Calculate all the eigenvalues of the Laplacian matrix, and then sort all the eigenvalues in ascending order to obtain the eigenvalue sequence;
[0075] S42. Select the first \(k\) eigenvalues in the eigenvalue sequence and calculate their corresponding eigenvectors, and use the \(k\) eigenvectors to form matrix \(U\) i , matrix \(U\) i constitutes the eigenvector space.
[0076] S5. Use the Discretize clustering algorithm to cluster the eigenvector space, and obtain the clustering result with the highest quality component as the component recognition result through the component quality function.
[0077] Specifically, step S5 includes:
[0078] S51. Use the classical algorithm to cluster matrix \(U\) i ; In order to select the classical algorithm, the present invention selects two classical algorithms, K-means and Discretize, to conduct experiments on four data sets respectively. The experimental results show that the component quality obtained by the two clustering algorithms is not very different, but there is a huge difference in the running time of the algorithms. K-means takes more time than the Discretize algorithm. Therefore, the Discretize algorithm is selected as the clustering algorithm.
[0079] The input of the clustering algorithm is the class call similarity matrix \(W\) and the number of clusters \(c\), and the output is the clustering result Com[]. The Discretize algorithm needs to input a number of clusters in advance, but the final number of clusters is determined by judging the clustering result through the component quality function. The algorithm process is shown in Table 1 below:
[0080] Table 1 Spectral Clustering
[0081]
[0082] S52. Use the component quality function ComQuality to evaluate the quality of clustering, and select the clustering result with the best quality among them.
[0083] Since the spectral clustering algorithm requires the number of clusters as an input parameter, in order to obtain the optimal solution of the clustering result, the SESc algorithm performs clustering for all possible numbers of clusters and selects the best clustering result among them. The evaluation process of the component quality function is that the quality of the clustering result is evaluated by the component quality function ComQuality. The possible range of the number of clusters is [2, N], where N is the number of classes in the software running event log SE. Select the clustering result that can obtain the highest quality of the identified components.
[0084] Specifically, the measurement of the clustering result of the present invention is mainly based on the following four component quality measurement indicators.
[0085] 1. The number of classes included in the component
[0086] To measure the quality of the identified components, the following measurements are first introduced:
[0087] The total number of components included in the software system (Number of component, abbreviated as NoC); the average number of classes included in each component (Average size number of component, abbreviated as AoC); the ratio of single class components (The ratio of single class components, abbreviated as RSC), that is, the proportion of components containing only one class among all components; the ratio of the largest component (The ratio of the largest component, abbreviated as RLC), indicating the proportion of the component containing the largest number of classes among all components; the ratio of intermediate components (The ratio of intermediate components, abbreviated as RIC), indicating the proportion of components of classes that are neither single class components nor the largest class components among all components.
[0088] Components with high RLC or high RSC cannot be regarded as good components. The ideal distribution is a normal distribution, and most components in the software system have a moderate and reasonable size (high RIC). Therefore, situations where there are relatively many single class components and the largest class components should be avoided as much as possible.
[0089] 2. Coupling degree
[0090] In a component-based software system, coupling represents the degree of tightness of interaction between one component and other components. The coupling degree between two components is calculated as the ratio of the actual value of the number of edges connecting them to the maximum number of edges that could connect all the nodes they contain.
[0091] Let G=(V, E) be a class interaction graph and CS be the set of identified components. For any components C 1 , C 2 ∈CS, their coupling degree can be calculated according to Equation (1):
[0092]
[0093] where: CouEdge = E ∩ ((C 1 × C 2 ) ∪ (C 2 × C 1 )) In a software system, the coupling degree of all components is calculated as shown in Equation (2):
[0094]
[0095] 3. Cohesion
[0096] In a component-based software system, cohesion represents the degree of tightness of class associations within the same component. The cohesion is calculated as the ratio of the actual number of edges connecting classes in a component to the number of edges that would connect all possible edges.
[0097] For any component C ∈ CS, its cohesion can be calculated according to Equation (3):
[0098]
[0099] where: CohEdge(C) = {(v 1 , v 2 ) ∈ E | v 1 , v 2 ∈ C}, represents a set of edges contained in component C, where v 1 , v 2 represent a class of the component respectively, and E represents the interaction relationship between classes. The cohesion of all components is calculated as shown in Equation (4):
[0100]
[0101] 4. Modular Quality
[0102] Cohesion and coupling metrics measure the quality of the identified components from two opposite perspectives. A better component-based software system should have high cohesion and low coupling. Modular Quality (MQ) combines the two. The higher the modular quality, the more likely the software is to have high cohesion and low coupling. The measurement method of modular quality is shown in Equation (5):
[0103] MQ(CS) = Cohesion(CS) - Coupling(CS) (5)
[0104] where MQ ∈ [-1, 1], and a higher MQ value usually means the software has a better architecture.
[0105] To obtain high-quality components, this paper defines the component quality function ComQuality. According to the above component quality metrics, the higher the RIC and MQ values of a component, the better the quality of the component. The component quality function ComQuality is shown in Equation (6), which is the average of RIC and MQ. In the SESC algorithm, the determination of the number of clusters is determined by this function. The algorithm traverses all possible numbers of clusters, calculates the component partitioning results obtained by clustering with different numbers of clusters, calculates the component quality according to the component quality function ComQuality, and selects the one with the highest quality as the final result.
[0106] The calculation process of the component quality function is as follows:
[0107] ComQuality = (MQ + RIC) / 2 (6)
[0108] MQ: represents modular quality, RIC: The ratio of intermediate components (abbreviated as RIC), which represents the proportion of components of classes that are neither single-class components nor the largest-class components among all components.
[0109] S6. Add the corresponding component attribute information (i.e., component name) to each group of software running cases in the software running event log according to the clustering result. A cluster in the clustering result corresponds to a component.
[0110] The input of the algorithm involved in the above process is the software running event log SE, and the output is the software running event log SE' with component attributes added. As shown in Table 2 below:
[0111] Table 2 Software Component Identification Algorithm SESC Based on Spectral Clustering
[0112]
[0113] The specific process of the algorithm is as follows:
[0114] The input of the algorithm is the software running event log SE, and the output is the software running event log SE' with component attributes added. First, according to the class call relationships in the software running event log SE, the algorithm calculates the class call similarity matrix W. It mainly uses the external method getClasscall() to extract the call relationships between classes from the software running event log SE and constructs the class call similarity matrix W. Then, it normalizes the values in the class call similarity matrix W using the external method normaliz(), as shown in steps 1 - 2 of the algorithm. Next, it determines the number of clusters. Using a traversal method, according to the number N of classes included in the software running event log SE, it performs spectral clustering on all possible numbers of clusters, as shown in steps 3 - 6 of the algorithm. Among them, the external method getClassNum() calculates the number N of all classes in the software running event log SE, and the possible number of clusters (i.e., the number of components) is [2, N]. At the same time, according to the result of each clustering, it calculates the quality of the components obtained by clustering according to the component quality function ComQuality. Among them, the external method SpectralClustering(W, i) calculates the clustering result when the number of clusters of the similarity matrix W is i, and the external method ComQuality(Com[]) calculates the quality of the components Com[] obtained according to the clustering result. Then, it selects the number of clustering clusters that can obtain the highest component quality, and the corresponding clustering result is used as the final component recognition result, as shown in steps 3 - 8 of the algorithm. The method SelectMax() selects the number of clustering clusters corresponding to the clustering result that can obtain the highest component quality. Finally, according to the clustering result, it adds component attribute information to each event in the software event log. Among them, the external method addComAttribute(SE, Com[]’) represents adding component information to the software running event log SE according to the clustering result Com[]’, as shown in steps 9 - 11 of the algorithm.
[0115] The present invention also provides a software component recognition system based on spectral clustering, including:
[0116] A data acquisition module, configured to acquire the software running event log SE = {c 0 , c 1 , …, c M} containing M groups of software running cases, and acquire the class set U Cl = {cl 1 , cl 2 , …, cl N} composed of all classes involved in the software running event log SE; where c m , m = 1, 2, …, M represents the m-th group of software running cases in the software running event log SE, and cl i , i = 1, 2, …, N represents the i-th class;
[0117] A class call similarity matrix construction module, which is used to construct a class call similarity matrix according to the software running event log SE and its class set U Cl Construct a class call similarity matrix;
[0118] An eigenvector space construction module, which is used to construct a degree matrix according to the class call similarity matrix, and calculate the Laplacian matrix based on the degree matrix; calculate all the eigenvalues of the Laplacian matrix and sort them in ascending order, and construct an eigenvector space through the eigenvectors of the first k eigenvalues;
[0119] A clustering output module, which is used to cluster the eigenvector space by using the Discretize clustering algorithm, and obtain the clustering result with the highest quality component as the component recognition result through the component quality function; after adding component attribute information to the software running event log SE through this component recognition result, output it.
[0120] In one embodiment, to verify the effectiveness of the method proposed by the present invention, the present invention conducted experiments. The algorithm SESc proposed by the present invention was compared with five other component recognition algorithms on four groups of public data sets from different aspects (as shown in Table 3).
[0121] The five algorithms are as follows: the first algorithm is the Newman spectral algorithm (denoted as NSA); the second algorithm is the moving node optimization algorithm of the Newman spectral algorithm (denoted as MVM); the third algorithm is the intelligent local movement algorithm (denoted as SLM); the fourth algorithm is the Louvain algorithm (denoted as LA); the fifth algorithm is the multi-level optimization algorithm of the Louvain algorithm (denoted as M-LA).
[0122] The public data sets are respectively: Lexi 0.1.1 1 、JGraphx 3.5.1 2 、JHotDraw 5.1 3 、Junit3.7 4 .
[0123] Table 1 Settings of comparison methods
[0124]
[0125] Figure 2It is the number of components identified by six algorithms for 4 datasets. It can be seen from the figure that the NoC and AoC values of the components identified by the SESC, SLM, LA, and M-LA algorithms are similar to the baseline values. The SESC algorithm achieves the same result as the baseline value on the Lexi 0.1.1 and JGraphx 3.5.1 datasets. The M-LA algorithm achieves the same result as the baseline value on the Lexi 0.1.1 and JHotDraw 5.1 datasets. However, the NoC value of the components identified by the NSA algorithm is higher than that of other algorithms, that is, the NSA algorithm identifies too many components for each software system. The NoC value of the components identified by the MVM algorithm is lower than that of other algorithms, that is, the MVM algorithm identifies fewer components. For the JUnit 3.7 dataset, compared with the six component identification algorithms, the baseline NoC value is too low, indicating that the organizational structure of this software package is not as good as that of other software packages. For the JGraphx 3.5.1 dataset, the number of components identified by the MVM, SLM, LA, and M-LA algorithms is small, while the NSA algorithm identifies more components. Only the SESC algorithm achieves the same result as the baseline value. In summary, the component identification result of the SESC algorithm is closer to the baseline value and is more accurate than other algorithms.
[0126] The NoC: The total number of components contained in the software system (Number of component, abbreviated as NoC);
[0127] The AoC: The average number of classes contained in each component (Average size number of component, abbreviated as AoC).
[0128] Figure 3 It is the size of the RSC and RIC values of the components identified by six algorithms for 4 datasets and the baseline. For the components identified for the same software system, the smaller the RSC value and the RIC value, the higher the quality of the identified components. It can be seen from the figure that the components identified by the SESC proposed in the present invention have achieved the best results on the four datasets.
[0129] The RSC: The ratio of single class components (abbreviated as RSC);
[0130] The RIC: The ratio of intermediate components (abbreviated as RIC).
[0131] Figure 4It is the ComQuality values of the components identified by six algorithms for 4 datasets and the baseline. As can be seen from the figure, both the ComQuality and RIC values of the components identified by the SESC algorithm reach the highest. In terms of the MQ value, although it does not perform higher than the NSA algorithm on some datasets, the RIC value of the NSA algorithm is lower, that is, the components identified do not have an advantage in terms of the number of classes. Generally speaking, the SESC algorithm still achieves the best results.
[0132] The MQ mentioned above: Modular Quality (abbreviated as MQ).
[0133] Figure 5 It is the time consumed by six algorithms to identify components for 4 datasets. The time consumption of the SESC algorithm includes the time to determine the number of clusters, that is, the time consumed to traverse all possible numbers of clusters and finally determine the optimal clustering result. As can be seen from the figure, except for the Lexi 0.1.1 data where the SESC algorithm consumes a longer time, the SESC algorithm consumes less time on the other three datasets and has an obvious advantage. The main reason is that the number of method calls in the Lexi 0.1.1 data is small, so the advantage of the SESC algorithm is not obvious, and other community detection algorithms consume less time for this data. However, for software operation data with a large number of method calls, the number of nodes in the class interaction graph increases and the graph structure becomes extremely complex. At this time, the time consumption of community detection algorithms including the LA algorithm increases greatly, while the SESC algorithm can still complete the identification of components in a short time.
[0134] The above experimental results show that the SESC algorithm proposed by the present invention can obtain components with the highest quality by traversing all numbers of clusters and determining the optimal clustering solution according to the component quality function, and the accuracy of the components identified by the algorithm can be verified by comparing with the baseline value. Finally, the SESC algorithm also has an advantage in time performance compared with the existing component identification algorithms.
[0135] In the present invention, unless otherwise clearly specified and defined, terms such as "installation", "setting", "connection", "fixation", "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal communication of two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0136] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A software component identification method based on spectral clustering, characterized in that: The following steps are involved: S1. Get the software running event log SE, and get all the classes involved in the software running event log SE to form a class set U Cl ; S2. Based on the software running event log SE and its class set U Cl Construct class call similarity matrix; Step S2: According to the software running event log SE and its class set U Cl The conversion results in a class call similarity matrix, including: S21. Treat each class as a node and connect them by edges based on the interaction between the classes; S22. Calculate the total number of interactions between every two classes, expressed as: Among them, W SE (cl i ,cl j ) represents class cl i With class cl j The total number of interactions between Represents class cl i With class cl j In the mth group of software running case c m The interaction results in Indicates that in the mth group of software running case c m Middle class cl i Called class cl j , Indicates that in the mth group of software running case c m Middle class cl j Called class cl i ; S23. The total number of interactions between every two classes is regarded as the edge weight of the corresponding edge, and an N×N class call similarity matrix W is constructed; S3. construct a degree matrix based on the class call similarity matrix, and calculate a Laplace matrix based on the degree matrix; The specific process of step S3 includes: S31. The calculation class calls the sum of the elements of each row of the similarity matrix and uses d i , i=1,2,…,N represents the sum of the elements in the i-th row of the class call similarity matrix; S32. Construct an N×N matrix D and use D ij represents the element in the i-th row and j-th column of matrix D; let D ii =d i And the rest of the elements are 0, and the degree matrix corresponding to the class call similarity matrix is obtained; S33. Calculate the Laplace matrix based on the degree matrix. The calculation expression is: It will be two. -1 (Two-W) Among them, Lap represents the Laplace matrix, De represents the degree matrix, and W represents the class call similarity matrix; S4. Calculate all eigenvalues of the Laplacian matrix and arrange them in ascending order, and construct an eigenvector space through the eigenvectors of the first k eigenvalues; S5. Use the Discretize clustering algorithm to cluster the feature vector space, and obtain the clustering result with the highest quality component as the component recognition result through the component quality function; S6. Add component attribute information to the software operation event log SE based on the component identification result.
2. The software component identification method based on spectral clustering according to claim 1, characterized in that: Step S1 obtains a software operation event log SE containing M groups of software operation cases = {c0, c1, ..., c M }, and obtain all classes involved in the software running event log SE to form a class set U Cl ={cl1,cl2,…,cl N }; where c m , m=1,2,…,M represents the mth group of software running cases in the software running event log SE, cl i , i=1,2,…,N represents the i-th class.
3. The software component identification method based on spectral clustering according to claim 1, characterized in that: Step S5 uses the Discretize clustering algorithm to cluster the feature vector space to obtain multiple clusters, evaluates the quality of each cluster through the component quality function, and selects the cluster with the best quality as the component identification result; the component quality function evaluation formula is: ComQuality=(MQ+RIC) / 2 Among them, MQ is the modular quality and RIC is the proportion of intermediate components.
4. The software component identification method based on spectral clustering according to claim 1, characterized in that: Step S6 includes: According to the clustering results, component attribute information is added to each group of software running cases in the software event log.
5. A software component identification system based on spectral clustering, characterized in that: include: The data acquisition module is used to acquire the software operation event log SE containing M groups of software operation cases = {c0, c1, ..., c M }, and obtain all classes involved in the software running event log SE to form a class set U Cl ={cl1,cl2,…,cl N }; where c m , m=1,2,…,M represents the mth group of software running cases in the software running event log SE, cl i , i=1,2,…,N represents the i-th class; The class call similarity matrix building module is used to construct the class call similarity matrix based on the software running event log SE and its class set U. Cl Construct class call similarity matrix; The class calls the Similarity Matrix building block, which includes: Step S2: According to the software running event log SE and its class set U Cl The conversion results in a class call similarity matrix, including: S21. Treat each class as a node and connect them by edges based on the interaction between the classes; S22. Calculate the total number of interactions between every two classes, expressed as: Among them, W SE (cl i ,cl j ) represents class cl i With class cl j The total number of interactions between Represents class cl i With class cl j In the mth group of software running case c m The interaction results in Indicates that in the mth group of software running case c m Middle class cl i Called class cl j , Indicates that in the mth group of software running case c m Middle class cl j Called class cl i ; S23. The total number of interactions between every two classes is regarded as the edge weight of the corresponding edge, and an N×N class call similarity matrix W is constructed; The eigenvector space construction module is used to construct a degree matrix based on the class call similarity matrix, and calculate the Laplace matrix based on the degree matrix; calculate all eigenvalues of the Laplace matrix and arrange them in ascending order, and construct the eigenvector space through the eigenvectors of the first k eigenvalues; The Laplace matrix is calculated based on the degree matrix, including: S31. The calculation class calls the sum of the elements of each row of the similarity matrix and uses d i , i=1,2,…,N represents the sum of the elements in the i-th row of the class call similarity matrix; S32. Construct an N×N matrix D and use D ij represents the element in the i-th row and j-th column of matrix D; let D ii =d i And the rest of the elements are 0, and the degree matrix corresponding to the class call similarity matrix is obtained; S33. Calculate the Laplace matrix based on the degree matrix. The calculation expression is: It will be two. -1 (Two-W) Among them, Lap represents the Laplace matrix, De represents the degree matrix, and W represents the class call similarity matrix; The clustering output module is used to cluster the feature vector space using the Discretize clustering algorithm, and obtain the clustering result with the highest quality component as the component identification result through the component quality function; the component attribute information is added to the software operation event log SE through the component identification result and then output.