Differentiation bipartite graph-based non-aligned multi-view clustering method, device and equipment
Through the method based on differentiated two-part graphs, anchor points are determined, initial two-part graphs are generated, and alignment is performed, the multi-view data clustering problem of misalignment of views is solved, and efficient and accurate clustering effect is achieved, providing convenience for web page data analysis and recommendation.
Patent Information
- Application Number
- CN202510297767.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to efficiently and accurately cluster multi-view data with misaligned views.
By determining the anchor points of each view in the multi-view dataset, the initial two-part graph is generated, the initial clustering is performed, the total contour coefficient is calculated, the most reliable view is determined, the alignment matrix is constructed, and the two-part graphs are aligned, and the multi-view dataset is finally guided to cluster.
It realizes efficient and accurate clustering of multi-view data with misaligned views, improves clustering speed and accuracy, and provides convenience for web page data analysis and recommendation.
Smart Images

Figure CN120147677A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a misaligned multi-view clustering method, device and equipment based on a differential bipartite graph. Background Art
[0002] Multi-view data refers to a data type in which the same object has multiple feature representations. For example, the same web page can be represented by the content on the web page and the links related to the web page. The web page content and the inter-webpage links constitute two views. The features on different views describe different aspects of the object, so the information contained in different views has consistency and complementarity. When traditional methods process multi-view data, it is required that the correspondence relationship between samples on different views is clear. However, in the multi-view data collected in practical applications, the correspondence relationship between samples on different views may be unknown and misaligned, forming misaligned multi-view data. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a misaligned multi-view clustering method, device and equipment based on a differential bipartite graph for efficiently and accurately clustering misaligned multi-view data.
[0004] The content of the present invention includes a misaligned multi-view clustering method based on a differential bipartite graph, including:
[0005] Determining the anchor points of each view in the multi-view dataset to be processed;
[0006] Generating an initial bipartite graph for a specific view based on the anchor points;
[0007] Performing initial clustering on the multi-view dataset;
[0008] Calculating the total silhouette coefficient of each view in the multi-view dataset based on the initial clustering result;
[0009] Determining the most reliable view based on the total silhouette coefficient of each view;
[0010] Constructing an alignment matrix for cross-view alignment;
[0011] Performing alignment processing on the bipartite graph based on the most reliable view and the alignment matrix;
[0012] Guiding the clustering of the multi-view dataset based on the aligned bipartite graph.
[0013] In one embodiment, the determining the anchor points of each view in the multi-view dataset to be processed includes:
[0014] Determining the types of each view in the multi-view dataset to be processed;
[0015] Perform non - negative processing on the features based on the same - type views to obtain a non - negative processed feature matrix;
[0016] Calculate the scores of each feature based on the non - negative processed feature matrix;
[0017] Select anchor points for each view based on the scores of the features.
[0018] In one embodiment, the selecting anchor points for each view based on the scores of the features includes:
[0019] Determine the initial anchor point based on the feature with the maximum score;
[0020] Normalize the scores of the features based on the initial anchor point;
[0021] Repeat the above two steps to determine the differential anchor points in each view.
[0022] In one embodiment, the generating an initial bipartite graph of a specific view based on the anchor points includes:
[0023] Generate an initial bipartite graph of a specific view based on the anchor points and the following formula:
[0024]
[0025] Where, and represent the elements of the \(i\) - th row and the \(j\) - th column of the \(i\) - th row of the bipartite graph \(B\) v The first - feature median and the second - feature median of the clustering samples in each view are calculated based on the initial clustering result. The first - feature median represents the average distance between the corresponding feature vector and other feature vectors belonging to the same initial clustering result, and the second - feature median represents the minimum average distance between the corresponding feature vector and feature vectors in other initial clustering results; represents the distance between the \(i\) - th sample and the \(j\) - th anchor point, \(\beta\) represents a parameter measuring the influence degree of the regularization term, \(n\) is the number of samples, and \(m\) is the number of anchor points.
[0026] In one embodiment, the calculating the silhouette coefficient of each view in the multi - view dataset based on the initial clustering result includes:
[0027] Calculate the silhouette coefficient of each feature in the view based on the first - feature median and the second - feature median;
[0028] Calculate and determine the total silhouette coefficient of the view based on the silhouette coefficients of each feature.
[0029] Calculate and determine the total silhouette coefficient of the view based on the silhouette coefficients of each feature.
[0030] In one embodiment, calculating the silhouette coefficient of each feature in the view based on the first feature median and the second feature median includes:
[0031]
[0032] where x i is a feature vector, a(x i ) is the first feature median, b(x i ) is the second feature median, and s(x i ) is the silhouette coefficient of the feature vector x i .
[0033] In one embodiment, determining the most reliable view based on the total silhouette coefficient of each view includes:
[0034] Determining the maximum value among the total silhouette coefficients of all the views;
[0035] Determining the view corresponding to the maximum value among the total silhouette coefficients as the most reliable view.
[0036] In one embodiment, aligning the bipartite graph based on the most reliable view and the alignment matrix includes:
[0037] Constructing a loss function;
[0038] Aligning the bipartite graph with the bipartite graph of the most reliable view based on the loss function and the alignment matrix.
[0039] Another embodiment of the present invention also provides a misaligned multi-view clustering device based on a differential bipartite graph, including:
[0040] A first determination module, configured to determine the anchor points of each view in the multi-view dataset to be processed;
[0041] A first generation module, configured to generate an initial bipartite graph of a specific view based on the anchor points;
[0042] An initial clustering module, configured to perform initial clustering on the multi-view dataset;
[0043] A first calculation module, configured to calculate the total silhouette coefficient of each view in the multi-view dataset based on the initial clustering result;
[0044] A second determination module, configured to determine the most reliable view based on the total silhouette coefficient of each view;
[0045] A first construction module, configured to construct an alignment matrix for cross-view alignment;
[0046] The first processing module is configured to perform alignment processing on the bipartite graph based on the most reliable view and the alignment matrix;
[0047] The clustering module is configured to guide the clustering of the multi-view data set based on the aligned bipartite graph.
[0048] Another embodiment of the present invention further provides an electronic device, including;
[0049] One or more processors;
[0050] A memory configured to store one or more programs;
[0051] When the one or more programs are executed by the one or more processors, the one or more processors implement the misaligned multi-view clustering method based on the differential bipartite graph as described in any one of the above.
[0052] The beneficial effects of the present invention include addressing the clustering problem of multi-view data with view misalignment. Different bipartite graphs are constructed for different views, allowing the anchor points of different views to be inconsistent. Additionally, by introducing a permutation matrix, the samples on the bipartite graphs of each view are aligned to utilize the consistent and complementary information of the multi-view data to complete efficient and accurate clustering of the multi-view data, improving the clustering speed and accuracy, facilitating web page data analysis and web page recommendation in daily operations, and increasing the analysis efficiency, analysis accuracy, as well as the recommendation efficiency and recommendation accuracy, etc.
[0053] Other features and advantages of the present application will be described in the subsequent description, and some of them will be obvious from the description or understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written description, claims, and drawings.
[0054] The technical solutions of the present application will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0055] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings required for the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a schematic flowchart of the misaligned multi-view clustering method based on the differential bipartite graph according to the embodiment of the present invention.
[0057] Figure 2Schematic flowchart of the misaligned multi-view clustering method based on a differential bipartite graph according to another embodiment of the present invention.
[0058] Figure 3 Schematic flowchart of the misaligned multi-view clustering method based on a differential bipartite graph according to another embodiment of the present invention.
[0059] Figure 4 Block diagram of the misaligned multi-view clustering device based on a differential bipartite graph according to an embodiment of the present invention. Detailed implementation manners
[0060] Next, specific embodiments of the present invention will be described in detail with reference to the accompanying drawings, but this is not a limitation to the present invention.
[0061] It should be understood that various modifications can be made to the embodiments disclosed herein. Therefore, the following description should not be regarded as limiting, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope of the present disclosure.
[0062] The accompanying drawings included in and constituting a part of this specification illustrate embodiments of the present disclosure, and together with the general description of the present disclosure given above and the detailed description of the embodiments given below are used to explain the principles of the present disclosure.
[0063] These and other features of the present invention will become apparent from the following description of the preferred forms of the embodiments given as non-limiting examples with reference to the accompanying drawings.
[0064] It should also be understood that although the present invention has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present invention, which have the features as described in the claims and thus are all within the protection scope defined thereby.
[0065] When combined with the accompanying drawings, the above and other aspects, features, and advantages of the present disclosure will become more apparent in view of the following detailed description.
[0066] Hereinafter, specific embodiments of the present disclosure will be described with reference to the accompanying drawings; however, it should be understood that the disclosed embodiments are only examples of the present disclosure, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details from obscuring the present disclosure. Therefore, the specific structural and functional details disclosed herein are not intended to be limiting, but only as a basis for the claims and a representative basis for teaching those skilled in the art to use the present disclosure in substantially any suitable detailed structure in various ways.
[0067] This specification may use phrases such as "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present disclosure.
[0068] Next, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0069] As Figure 1 shown, an embodiment of the present invention provides a misaligned multi-view clustering method based on a differential bipartite graph, including:
[0070] S1: Determine the anchor points of each view in the multi-view dataset to be processed;
[0071] S2: Generate an initial bipartite graph for a specific view based on the anchor points;
[0072] S3: Perform initial clustering on the multi-view dataset;
[0073] S4: Calculate the total silhouette coefficient of each view in the multi-view dataset based on the initial clustering result;
[0074] S5: Determine the most reliable view based on the total silhouette coefficient of each view;
[0075] S6: Construct an alignment matrix for cross-view alignment;
[0076] S7: Perform alignment processing on the bipartite graph based on the most reliable view and the alignment matrix;
[0077] S8: Guide the clustering of the multi-view dataset based on the aligned bipartite graph.
[0078] The method of this embodiment can be applied to the clustering analysis of multi-view datasets such as web pages. For example, by analyzing the web pages commonly used by users, it can determine user preferences and complete user web page recommendations. Through the method of this embodiment, the efficiency and accuracy of clustering analysis can be improved, and reliable basis can be provided for subsequent applications such as determining user preferences and making recommendations for users, thereby enhancing the recommendation efficiency and accuracy.
[0079] Based on the above, it can be seen that the method described in this embodiment is actually aimed at the clustering problem of multi-view data with misaligned views. Different bipartite graphs are constructed for different views, allowing the anchor points of different views to be inconsistent. In addition, by introducing a permutation matrix, the samples on the bipartite graphs of each view are aligned, so as to utilize the consistent and complementary information of multi-view data to complete the efficient and accurate clustering of multi-view data, improve the clustering speed and accuracy, facilitate web page data analysis and web page recommendation in daily life, and increase the analysis efficiency, analysis accuracy, as well as recommendation efficiency and recommendation accuracy, etc.
[0080] In one embodiment, assume a multi-view data set consisting of n samples on n views, and the feature matrix of the v-th view is V where d is the feature dimension of the v-th view, υ = 1, 2, …, n v , and the number of clusters of the data set is c. There are m anchor points on each view, and the feature matrix of the anchor points of the v-th view can be expressed as V is the set of all anchor points of the data set. The bipartite graph on a specific view is the affinity matrix of the feature matrix X v and the feature matrix Y υ of the anchor points, or the weight matrix. The element v in the bipartite graph B represents the weight of the edge connecting the i-th instance and the j-th anchor point. Thus, the initial bipartite graph of the multi-view data set is obtained
[0081] Specifically, determining the anchor points of each view in the multi-view data set to be processed includes:
[0082] S9: Determine the types of each view in the multi-view data set to be processed;
[0083] S10: Perform non-negative processing on the features based on the views of the same type to obtain a non-negatively processed feature matrix;
[0084] S11: Calculate the scores of each feature based on the non-negatively processed feature matrix;
[0085] S12: Select anchor points for each view based on the scores of the features.
[0086] Among them, as Figure 2 shown, selecting anchor points for each view based on the scores of the features includes:
[0087] S13: Determine the initial anchor points based on the features with the maximum scores;
[0088] S14: Normalize the scores of the features based on the initial anchor points;
[0089] S15: Repeat the above two steps to determine the differentiated anchor points in each view.
[0090] For example, in this embodiment, the direct alternating sampling DAS method is used. Considering that the selected anchor points should be able to effectively cover the entire data point cloud, and the features of similar samples should be similar, and the features of different samples should be different, alternating sampling is performed according to the sample features, which has good performance and stability. The specific method for selecting anchor points includes:
[0091] For a specific view feature matrix composed of n samples with a feature dimension of d v First, based on the feature similarity of samples of the same class, the features are non-negatively processed. Let the minimum value in the feature vector be
[0092]
[0093] to obtain the non-negatively processed feature matrix After that, calculate the scores s = [s 1 , s 2 , …, s n of all samples T .
[0094]
[0095] When the samples belong to the same cluster, the scores will be very close; when they belong to different clusters, the score differences are large. Select the anchor points one by one according to the score size: select the one with the largest score as the initial anchor point:
[0096] Subsequently, normalize the scores and update them:
[0097]
[0098] s i = s i (1 - s i )
[0099] Repeat the above process m times to obtain m non-repeating anchor points. Search for anchor points on each view, and the feature vectors of m differentiated anchor points can be obtained on each view to form the feature matrix of the anchor points
[0100] captures the view complementary information of multi-view data with view misalignment.
[0101] In another embodiment, generating the initial bipartite graph of a specific view based on the anchor points includes:
[0102] S16: Generate the initial bipartite graph of a specific view based on the anchor points and the following formula:
[0103]
[0104] where and represent the bipartite graph B v The element in the i-th row and the j-th column of the i-th row, represents the distance between the i-th sample and the j-th anchor point, β represents a parameter measuring the influence degree of the regularization term, n is the number of samples, and m is the number of anchor points.
[0105] Specifically, after obtaining the differentiated anchor points on all views, an initial bipartite graph for a specific view is generated. On a specific view, we learn a non-negative and normalized initial bipartite graph B v . Considering that the edge weight between samples and anchor points with smaller distances should be larger, a bipartite graph is constructed by the following formula:
[0106] where, and represent the element in the i-th row and the j-th column of the i-th row of the bipartite graph B v , represents the distance between the i-th sample and the j-th anchor point, β represents a parameter measuring the influence degree of the regularization term, n is the number of samples, and m is the number of anchor points. The first term of the above formula is used to learn the structure of the bipartite graph, making the connection between the sample features and the anchor points closer to it strong, and the connection with the anchor points farther from it weak; the second term plays a role in regularization and is used to control the sparsity of the bipartite graph. Based on this, the above formula has the following closed-form solution:
[0107]
[0108] where (v) + = max(0, v), when j > k, this solution is zero. Therefore, each sample feature is only connected to the k nearest anchor points. When obtaining the initial bipartite graph of a specific view through the above method, the bipartite graph is quite sparse, and only the important terms in the graph need to be optimized, which can greatly reduce the calculation time. Each sample is only connected to n V × k connections among the m anchor points. V
[0109] Furthermore, calculating the silhouette coefficient of each view in the multi-view dataset based on the initial clustering result includes:
[0110] S17: Calculating the first feature median and the second feature median of the clustered samples in each view based on the initial clustering result. The first feature median represents the average distance between the corresponding feature vector and other feature vectors belonging to the same initial clustering result, and the second feature median represents the minimum average distance between the corresponding feature vector and the feature vectors in other initial clustering results;
[0111] S18: Calculating the silhouette coefficient of each feature in the view based on the first feature median and the second feature median;
[0112] S19: Determine the total silhouette coefficient of the view based on the silhouette coefficients calculated for each of the said features.
[0113] Exemplarily, the silhouette coefficient can measure the degree of internal clustering tightness and the degree of clustering dispersion, and measures the within-cluster distance and the between-cluster distance. The data set in this embodiment has been pre-clustered by a clustering method, such as the k-means method, that is, its clustering result, which is the initial clustering result, has been obtained. According to the initial clustering result, this embodiment can calculate two intermediate values corresponding to the sample features on a specific view, namely the first feature intermediate value and the second feature intermediate value:
[0114]
[0115] Represents the Euclidean distance between the feature vectors x i and x j |C a | and |C b | represent the number of samples in the clusters C a and C b Intuitively understood, a(x i ) represents the average distance between the feature vector x i and other feature vectors belonging to the same initial clustering result, and b(x i ) represents the minimum average distance between the feature vector and the feature vectors in other initial clustering results.
[0116] Subsequently, the silhouette coefficient s of the feature vector x i composed of a(x i ) and b(x i ) can be obtained:
[0117]
[0118] The above formula can be abbreviated into a formula, that is
[0119]
[0120] The said x i is the feature vector, the said a(x i ) is the first feature intermediate value, the said b(x i ) is the second feature intermediate value, and the said s(x i ) is the silhouette coefficient of the feature vector x i .
[0121] Calculating and taking the average of all feature vectors of a cluster can measure the degree of aggregation of this cluster and the degree of separation from other clusters. And taking the average of all clusters, that is, all feature vectors, can obtain the clustering quality of the entire data set. Thus, there is the silhouette coefficient of a specific view:
[0122]
[0123] As can be seen from the above formula, S v The larger it is, the greater the inter-class distance and the smaller the intra-class distance, that is, the easier it is to distinguish the clusters; on the contrary, the smaller it is, the smaller the inter-class distance and the larger the intra-class distance, that is, the easier it is for the clusters to be confused.
[0124] Furthermore, as Figure 3 shown, determining the most reliable view based on the silhouette coefficient of each view includes:
[0125] S20: Determine the maximum value among the silhouette coefficients of all the views;
[0126] S21: Determine the view corresponding to the maximum value among the silhouette coefficients as the most reliable view.
[0127] For example, determine the silhouette coefficient S v of each view, and take the view with the largest S v as the most reliable view ν 0 :
[0128] After obtaining the most reliable view, bipartite graph alignment can be performed. Consider an alignment matrix ij where all elements p are either 0 or 1, and each row and each column has only one 1, that is, it is a row permutation matrix. Therefore, there is: Among them, is the feature matrix after alignment adjustment, which is the aligned multi-view data set.
[0129] Since the bipartite graph is the affinity matrix of samples and anchors, the selection of anchors has nothing to do with whether the data set is aligned or not, and only relates to the properties of the feature vectors. Therefore, the alignment matrix acting on the feature matrix and the bipartite graph is equivalent. Considering the initial bipartite graph of unaligned data, it can be inferred that there is an alignment matrix P ν for a specific view, satisfying Among them, is the bipartite graph after alignment adjustment, which is the aligned multi- Figure 2 view data set.
[0130] Since the most reliable view ν has been obtained0 , considering that there are certain common characteristics in the bipartite graphs on different views, therefore, on a specific view, the alignment matrix P can be learned ν , and the originally misaligned bipartite graph B ν is aligned with the bipartite graph of the most reliable view . Specifically, the alignment process of the bipartite graph based on the most reliable view and the alignment matrix includes:
[0131] S22: Construct a loss function;
[0132] S23: Align the bipartite graph with the bipartite graph of the most reliable view based on the loss function and the alignment matrix.
[0133] The loss function proposed in this embodiment can be expressed as follows:
[0134]
[0135] S1 nVm =n V 1 n , S≥0
[0136] where is the alignment matrix, is the initial bipartite graph on a specific view, is the initial bipartite graph on the most reliable view, is the fused aligned multi-view Figure 2 bipartite graph set:
[0137]
[0138] is the normalized Laplacian matrix:
[0139]
[0140] is the degree matrix of Z, is the diagonal matrix established by S according to the result of its column sum, can be represented in a block structure as:
[0141]
[0142] where α is a hyperparameter that measures the proportion of the cross-view alignment term. At this time, all terms of the loss function are obtained, and then optimizing it can be used for solving.
[0143] During the optimization process, since there is a rank constraint in the loss function but the non-convex rank constraint is difficult to solve, so the corresponding terms of the loss function are solved to meet the rank constraint condition. Let The eigenvalues of Since the normalized Laplacian matrix is clearly positive semi - definite, we have Then the condition: the multiplicity of the eigenvalue 0 of the normalized Laplacian matrix equals c can be transformed into its smallest c eigenvalues being 0. Let the i - th smallest eigenvalue of When the rank constraint is satisfied. Thus, the objective function can be rewritten as
[0144]
[0145] s.t.
[0146]
[0147] where λ is a parameter to ensure that the eigenvalues of the Laplacian matrix are small enough. When λ is large enough, will tend to 0 infinitely, which guarantees that the smallest c eigenvalues are 0. Considering that hyperparameters are difficult to adjust, the parameter here does not need to be adjusted manually, but is automatically adjusted based on the number of 0 eigenvalues obtained. First, initialize λ to a small value to prevent too many connected components. In the paper, it is set as λ 0 = 0.1. In each iteration, update S to get a new After that, adjust the value of according to the size relationship between and c:
[0148]
[0149] λ new represents the new λ, and automatically update λ until the rank constraint is satisfied, that is Stop updating the iteration of S and jump out of updating the alignment matrix.
[0150] Next, use the Ky Fan theorem to transform the calculation of eigenvalues into the trace of matrix products. So we have:
[0151]
[0152] s.t.
[0153]
[0154] F T F = I
[0155] Now, there are three variables to be optimized: P, S, and F. Next, we use the method of updating one variable and fixing the other two variables to optimize them in turn.
[0156] 1. Fix S and F, and optimize P
[0157] When S and F are fixed, the loss function can be simplified to
[0158]
[0159] s.t.P ν 1 n = 1 n
[0160]
[0161] Converting the calculation of the F-norm of the matrix to the calculation of the trace, the above equation is equivalent to
[0162]
[0163] Since is the fused aligned multi-view Figure 2 bipartite graph set obtained by row-wise splicing, and the trace of the splicing matrix is equal to the sum of the traces of its blocks. Therefore, equation (24) can be simplified to minimizing the trace on a specific view
[0164]
[0165] where S v represents the block of the learned fused bipartite graph S by view, and the above equation holds for any view ν. Since constrained optimization is rather troublesome, we can discard the constraint conditions during the solution on the premise that the impact on the result is small. The above equation can be simply optimized by taking the derivative and setting it to 0, so the optimization of P is equivalent to
[0166]
[0167] Since B v B vT is obviously non-invertible and P v cannot be directly obtained. Therefore, adding the term to the loss function is equivalent to adding a regularization term, which does not affect the optimization result of the loss function. The result of taking the derivative of it on a specific view is δP ν , and the above equation is converted to
[0168]
[0169] where is invertible, and then we can get
[0170]
[0171] When the dataset is large, the matrix inversion calculation amount is extremely large, and the complexity is Therefore, use Sherman-Morrison to transform the inverse of the matrix:
[0172]
[0173] The complexity of the inverse part is reduced to Since the number of anchor points is much smaller than the number of samples m << n, the amount of calculation can be greatly reduced. So far, the optimization of P is completed.
[0174] 2. Fix P and S, and optimize F
[0175] When P and S are fixed, the loss function can be simplified to
[0176]
[0177] F T F = I
[0178] Since The computational complexity is According to 's definition, Z is the augmented graph of S, The subtraction part of can also be regarded as is a matrix, whose column f 1 ; f 2 ;...; is 's smallest c eigenvectors, which can be clearly obtained from 's block. Obviously, the first n V m rows of the eigenvector f represent the indices of the anchor points (multiplying with each column of the bipartite graph S, i.e., the anchor points), and the last n rows represent the indices of the samples (multiplying with each row of the bipartite graph S, i.e., the samples). Therefore, F can be divided into
[0179]
[0180] where The constraint condition is
[0181] At this time, expanding directly gives
[0182]
[0183] Obviously, the sum of the first term and the fourth term of the above formula is the identity matrix, and the middle two terms are transposes of each other. Therefore, there is no need to calculate the augmented graph Z, and only the bipartite graph S needs to be calculated. Thus, the above formula can be simplified to:
[0184]
[0185] where, Since the normalized Laplacian matrix is converted to the Laplacian matrix, finding the minimum is equivalent to finding the maximum. The computational complexity of the above equation is
[0186] By performing SVD to solve it.
[0187] In the SVD decomposition, for the problem The trace of the product is:
[0188]
[0189] A T A + B T B = I
[0190] The optimal solution of the above equation is and U is the left singular vector of X, and V is the right singular vector of X, corresponding to the largest c singular values of X.
[0191] Substituting the above equation into the SVD decomposition, we have A = F( n ), Let its left singular vector be U and the right singular vector be V, then F can be expressed as
[0192]
[0193] So far, the optimization of F is completed.
[0194] 3. Fix P and F, and optimize S
[0195] When P and F are fixed, the loss function can be simplified to
[0196]
[0197] For the term It has been expanded when optimizing F. Now expand the expansion result by rows:
[0198]
[0199] where D = diag(S T 1 n ), and there is is the degree matrix of Z
[0200]
[0201] f i represents the i-th row of F, represents the j-th row of The i-th row of F (n) where z ij and s ij represent the elements corresponding to the i-th row and j-th column of Z and S respectively.
[0202] Expanding the first term of the above formula by row, its simplified form can be obtained:
[0203]
[0204] where is the element corresponding to the i-th row and j-th column of the initial joint bipartite graph matrix Since the minimization of s ij is independent of each other, we can optimize S by row. Merging the above formula by row
[0205]
[0206] where and s i represent respectively and the i-th row of S, represents the result of merging the elements by row. The above formula can obtain a closed-form solution through the following algorithm.
[0207] For the constrained minimization problem
[0208]
[0209] it can be solved using the Lagrange multiplier method, and its Lagrangian function is
[0210]
[0211] There is It can be obtained by substituting into the function and finding the root of by Newton's method. Thus, the optimal solution α * is obtained.
[0212] So far, s i is obtained, and the optimization of S is completed.
[0213] As Figure 4 shown, another embodiment of the present invention also provides a misaligned multi-view clustering device 100 based on a differential bipartite graph, including:
[0214] A first determination module, configured to determine the anchor points of each view in the multi-view dataset to be processed;
[0215] A first generation module, configured to generate an initial bipartite graph for a specific view based on the anchor points;
[0216] An initial clustering module for performing initial clustering on the multi-view data set;
[0217] A first calculation module for calculating the silhouette coefficient of each view in the multi-view data set based on the initial clustering result;
[0218] A second determination module for determining the most reliable view based on the silhouette coefficient of each view;
[0219] A first construction module for constructing an alignment matrix for cross-view alignment;
[0220] A first processing module for performing alignment processing on the bipartite graph based on the most reliable view and the alignment matrix;
[0221] A clustering module for guiding the clustering of the multi-view data set based on the aligned bipartite graph.
[0222] In one embodiment, determining the anchor points of each view in the multi-view data set to be processed includes:
[0223] Determining the types of each view in the multi-view data set to be processed;
[0224] Performing non-negative processing on the features based on the views of the same type to obtain a non-negative processed feature matrix;
[0225] Calculating the scores of each feature based on the non-negative processed feature matrix;
[0226] Selecting anchor points for each view based on the scores of the features.
[0227] In one embodiment, the selecting anchor points for each view based on the scores of the features includes:
[0228] Determining an initial anchor point based on the feature with the maximum score;
[0229] Normalizing the scores of the features based on the initial anchor point;
[0230] Repeatedly execute the above two steps to determine the differential anchor points in each view.
[0231] In one embodiment, generating an initial bipartite graph of a specific view based on the anchor points includes:
[0232] Generating an initial bipartite graph of a specific view based on the anchor points and the following formula:
[0233]
[0234] Wherein, and Represents the bipartite graph B v The element in the i-th row and the j-th column of the i-th row, Represents the distance between the i-th sample and the j-th anchor point. β represents a parameter for measuring the influence degree of the regularization term. n is the number of samples, and m is the number of anchor points.
[0235] In one embodiment, calculating the silhouette coefficient of each view in the multi-view dataset based on the initial clustering result includes:
[0236] Calculating a first feature median value and a second feature median value of the clustered samples in each view based on the initial clustering result. The first feature median value represents the average distance between the corresponding feature vector and other feature vectors belonging to the same initial clustering result, and the second feature median value represents the minimum average distance between the corresponding feature vector and feature vectors in other initial clustering results;
[0237] Calculating the silhouette coefficient of each feature in the view based on the first feature median value and the second feature median value;
[0238] Calculating and determining the silhouette coefficient of the view based on the silhouette coefficients of each feature.
[0239] In one embodiment, calculating the silhouette coefficient of each feature in the view based on the first feature median value and the second feature median value includes:
[0240]
[0241] The x i Is a feature vector, the a(x i ) is the first feature median value, the b(x i ) is the second feature median value, and the s(x i ) is the silhouette coefficient of the feature vector x i .
[0242] In one embodiment, determining the most reliable view based on the silhouette coefficient of each view includes:
[0243] Determining the maximum value among the silhouette coefficients of all the views;
[0244] Determining the view corresponding to the maximum value of the silhouette coefficient as the most reliable view.
[0245] In one embodiment, aligning the bipartite graph based on the most reliable view and the alignment matrix includes:
[0246] Constructing a loss function;
[0247] Based on the loss function, the alignment matrix is used to align the bipartite graph with the bipartite graph of the most reliable view.
[0248] Another embodiment of the present invention further provides an electronic device, including:
[0249] One or more processors;
[0250] A memory configured to store one or more programs;
[0251] When the one or more programs are executed by the one or more processors, the one or more processors implement the misaligned multi-view clustering method based on the differential bipartite graph as described in any of the above.
[0252] Furthermore, an embodiment of the present invention further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the misaligned multi-view clustering method based on the differential bipartite graph as described above. It should be understood that each solution in this embodiment has the corresponding technical effects in the above method embodiment, and will not be elaborated here.
[0253] Furthermore, an embodiment of the present invention also provides a computer program product, the computer program product is tangibly stored on a computer-readable medium and includes computer-readable instructions, and the computer-executable instructions, when executed, cause at least one processor to execute the misaligned multi-view clustering method based on the differential bipartite graph in the above embodiments.
[0254] It should be noted that the computer storage medium of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable medium can, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access storage medium (RAM), a read-only storage medium (ROM), an erasable programmable read-only storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only storage medium (CD-ROM), an optical storage medium, a magnetic storage medium, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program configured to be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, antenna, optical cable, RF, etc., or any suitable combination of the above.
[0255] In addition, those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0256] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1one or more processes and / or blocks Figure 1 a system for the functions specified in one or more blocks
[0257] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction system that implements the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 one or more blocks
[0258] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; under the concept of this application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of this application as described above, and they are not provided in detail for the sake of brevity.
Claims
1. A non-aligned multi-view clustering method based on differentiated bipartite graphs, characterized in that: include: Determine the anchor point of each view in the multi-view dataset to be processed; generating an initial bipartite graph of a specific view based on the anchor point; performing initial clustering on the multi-view dataset; Calculating a total silhouette coefficient of each view in the multi-view dataset based on the initial clustering result; determining the most reliable view based on the total silhouette coefficient of each view; Construct an alignment matrix for cross-view alignment; Performing alignment processing on the bipartite graph based on the most reliable view and the alignment matrix; The multi-view dataset is clustered based on the aligned bipartite graph.
2. The unaligned multi-view clustering method based on differentiated bipartite graph according to claim 1, characterized in that: The determining of the anchor point of each view in the multi-view data set to be processed includes: Determining the type of each view in the multi-view data set to be processed; Performing non-negative processing on features based on the views of the same type to obtain a feature matrix after non-negative processing; Calculate the score of each feature based on the feature matrix after non-negative processing; Anchor points are selected for each view based on the scores of the features.
3. The unaligned multi-view clustering method based on differentiated bipartite graph according to claim 2, characterized in that: The selecting of anchor points for each view based on the score of the feature includes: Determine the initial anchor point based on the feature with the maximum score; Normalizing the score of the feature based on the initial anchor point; Repeat the above two steps to determine the differentiated anchor points in each view.
4. The unaligned multi-view clustering method based on differentiated bipartite graph according to claim 1, characterized in that: The generating an initial bipartite graph of a specific view based on the anchor point comprises: Generate an initial bipartite graph for a specific view based on the anchor point and the following formula: in, and Represents a bipartite graph B v The element in the i-th row and the i-th row and j-th column of represents the distance between the i-th sample and the j-th anchor point, β represents the parameter that measures the influence of the regularization term, n is the number of samples, and m is the number of anchor points.
5. The unaligned multi-view clustering method based on differentiated bipartite graph according to claim 1, characterized in that: The calculating the total silhouette coefficient of each view in the multi-view dataset based on the initial clustering result comprises: Calculating a first feature middle value and a second feature middle value of cluster samples in each view based on the initial clustering result, wherein the first feature middle value represents an average distance between a corresponding feature vector and other feature vectors in the same initial clustering result, and the second feature middle value represents a minimum average distance between the corresponding feature vector and feature vectors in other initial clustering results; Calculating a silhouette coefficient of each feature in the view based on the first feature intermediate value and the second feature intermediate value; An overall silhouette coefficient of the view is determined based on the silhouette coefficient calculations of the features.
6. The method for clustering unaligned multi-views based on differentiated bipartite graphs according to claim 5, characterized in that: The calculating the silhouette coefficient of each feature in the view based on the first feature intermediate value and the second feature intermediate value includes: The x i is a eigenvector, the a(x i ) is the first characteristic intermediate value, and b(x i ) is the second characteristic intermediate value, the s(x i ) is the feature vector x i The silhouette coefficient.
7. The unaligned multi-view clustering method based on differentiated bipartite graph according to claim 1, characterized in that: The determining the most reliable view based on the total silhouette coefficient of each view comprises: Determining the maximum value of the total silhouette coefficients of all the views; The view corresponding to the maximum value in the total silhouette coefficient is determined as the most reliable view.
8. The unaligned multi-view clustering method based on differentiated bipartite graph according to claim 1, characterized in that: The aligning the bipartite graph based on the most reliable view and the alignment matrix includes: Construct loss function; The bipartite graph is aligned with the bipartite graph of the most reliable view based on the loss function and the alignment matrix.
9. A non-aligned multi-view clustering device based on differentiated bipartite graph, characterized in that: include: A first determination module, configured to determine an anchor point of each view in a multi-view data set to be processed; A first generating module, configured to generate an initial bipartite graph of a specific view based on the anchor point; An initial clustering module, used for performing initial clustering on the multi-view dataset; A first calculation module, configured to calculate a total silhouette coefficient of each view in the multi-view dataset based on an initial clustering result; A second determination module, configured to determine the most reliable view based on the total silhouette coefficient of each view; A first building module, for building an alignment matrix for cross-view alignment; A first processing module, configured to perform alignment processing on the bipartite graph based on the most reliable view and the alignment matrix; A clustering module is used to guide the multi-view dataset to cluster based on the aligned bipartite graph.
10. An electronic device, characterized in that: include; one or more processors; a memory configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the unaligned multi-view clustering method based on a differentiated bipartite graph as described in any one of claims 1 to 8.