Incomplete Multi-View Clustering Method and System Based on Local Structure and Balance Perception

By designing an incomplete multi-view consistent clustering representation learning model based on local structure and balance perception, the existing methods solve the problems of high computational complexity, high memory consumption and unstable clustering performance when dealing with large-scale incomplete multi-view data, and achieve efficient and stable clustering effect.

CN115311483BActive Publication Date: 2025-05-27HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210979979.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2025-05-27
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

The existing incomplete multi-view clustering method has high computational complexity, high memory consumption, and unstable clustering performance when processing large-scale incomplete multi-view data.

Method used

A learning model for incomplete multi-view consistent clustering representation based on local structure and equilibrium perception is designed, and the unique clustering result is obtained by learning consistent representations with probability characteristics between views. The model integrates geometric structure-keeping and consistent characterization learning into a concise model, avoiding additional constraints and penalties parameters, and introducing balanced-aware learning techniques to avoid over-dividing of samples.

Benefits of technology

It realizes efficient and stable clustering performance, reduces computational complexity and memory consumption, and is suitable for clustering tasks with large-scale incomplete multi-view data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311483B_ABST
    Figure CN115311483B_ABST
Patent Text Reader

Abstract

The present invention discloses an incomplete multi-view clustering method and system based on local structure and balance perception, including: for the clustering task of incomplete multi-view data, designing an incomplete multi-view consistent clustering representation learning model with probability characteristics based on local structure and balance perception; preprocessing the incomplete multi-view data with a given view missing prior position index matrix; according to the preprocessed data, designing an alternating iteration optimization-based method to solve the variables based on the variables contained in the incomplete multi-view consistent clustering representation learning model, so as to achieve the purpose of model optimization, and obtaining the clustering results of all samples by using the optimal shared consistent representation matrix obtained after optimization. The model designed by the method of the present invention is an incomplete multi-view clustering model with interpretability, high efficiency and stable clustering results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of machine learning and pattern recognition, and particularly relates to an incomplete multi-view clustering method and system based on local structure and balance perception. Background Art

[0002] In the past few years, a large amount of multi-view data collected by different sensors or in different ways has emerged in different fields or industries, and the need for multi-view data clustering (MVC) has also emerged in various applications. For example, in order to predict the possible development trend of Alzheimer's disease, a multi-view clustering model of consistent representation learning has been proposed, and each sample of the model is represented by two types of brain magnetic resonance imaging data; in addition, the multi-view clustering model based on non-negative matrix factorization (NMF) has also achieved good results in web project recommendation. Generally speaking, traditional MVC methods are all based on the complete view assumption, that is, all samples can completely observe their complete view feature information. When using traditional multi-view clustering methods to process incomplete multi-view data clustering tasks, samples with missing views must be removed in advance, and only samples with complete views can be clustered. In fact, in many practical applications, such as recommendation systems and Alzheimer's disease diagnosis, the actual data collected is often incomplete data with missing views. Therefore, the research on incomplete multi-view clustering is of great significance, and it has more promotion value and application value than traditional multi-view clustering methods.

[0003] In recent years, scholars at home and abroad have also successively proposed some incomplete multi-view clustering models. For example, the canonical correlation analysis strategy is introduced into incomplete two-view clustering, and the missing information of another incomplete kernel matrix can be complemented by using the complete kernel matrix of a complete view. An incomplete multi-kernel k-means method with kernel complementation characteristics (IMKKM-MKC) has also been proposed to solve the incomplete multi-view clustering problem in the case of missing views of any view. In addition to these kernel complementation-based methods, an incomplete multi-view clustering method based on graph complementation designs a model from the perspective of graph learning, which can recover the missing information of multiple incomplete graph matrices and generate a consistent representation of the views. A unified embedding alignment framework (UEAF) based on matrix factorization designs a joint optimization model that can simultaneously recover missing view information and learn consistent representations. The above methods generally adopt the idea of restoring missing information to solve the incomplete multi-view clustering problem.

[0004] Although the method of missing information completion can, to some extent, solve the problem of incomplete multi-view clustering under partial view information missing, such methods still have the following disadvantages: 1) Almost all methods divide the clustering task into two independent and unrelated stages, that is, first perform graph or representation learning, and then implement k-means or spectral clustering to obtain the final incomplete multi-view clustering results. On the one hand, these methods cannot guarantee that the obtained graph or representation is a clustering-friendly representation that can achieve the best clustering performance; on the other hand, since these methods all use k-means to generate the final clustering results, and running k-means k times will produce k different clustering results, so these methods cannot obtain a stable and unique clustering solution. 2) These methods generally have high computational complexity and memory consumption, resulting in being unsuitable for handling the clustering tasks of "large-scale" incomplete multi-view data.

[0005] In summary, in recent years, to solve the challenging problem of multi-view data clustering with missing views, many incomplete multi-view clustering methods have been proposed. However, most of the existing methods are not suitable for the clustering tasks of large-scale incomplete multi-view data, and the clustering performance is unstable. Summary of the Invention

[0006] In view of the above problems, the present invention provides an incomplete multi-view clustering method and system based on local structure and balance perception. For the efficient learning problem of incomplete multi-view data, an incomplete multi-view Figure 1 consistent clustering representation learning model is designed, and this model obtains a unique clustering result by learning the consistent representation with probability characteristics between views.

[0007] In the first aspect of the present invention, an incomplete multi-view clustering method based on local structure and balance perception includes the following steps:

[0008] Establish a model: For the clustering task of incomplete multi-view data, design an incomplete multi-view Figure 1 consistent clustering representation learning model with probability characteristics based on local structure and balance perception. The specific model is:

[0009]

[0010]

[0011] Among them, represents the basis matrix of the v-th perspective, m v represents the feature dimension of the v-th view, d represents the dimension of the consistent representation space, P ∈ R d×n represents the shared consistent representation matrix of incomplete multi-view data, n represents the total number of samples of incomplete multi-view data, α = [α 1,...,α l is a learnable weight vector, 1∈R d denotes a d-dimensional column vector with all element values being 1, α v denotes the v-th element in the vector α, r is a positive integer not less than 2, denotes the element α in the vector α v to the r-th power, λ is the penalty term parameter, l represents the number of views, n v denotes the number of samples without missing values in the v-th view, I is the identity matrix, I i,j denotes the element value at the (i, j)-th row and column position of the identity matrix, denotes the similarity relationship between the i-th sample and the v-th view of the j-th sample, denotes the matrix X (v) of the i-th column vector, denotes the matrix set composed of samples without missing values in the v-th view, denotes the matrix G (v) of the j-th column vector, is a binary matrix of 0 and 1;

[0012] Data preprocessing: For the incomplete multi-view data of the given view missing prior position index matrix Z perform preprocessing;

[0013] Optimization model: According to the preprocessed data and the incomplete multi-view Figure 1 induced clustering representation learning model, for the variables P, α and the introduced auxiliary variables Q, Lagrange multipliers C and positive penalty parameter μ, design an alternating iteration optimization-based method to solve the variables to achieve the purpose of model optimization, where:

[0014] Solve for U (v) of the optimization problem: Obtain the variable U (v) The optimal solution of is U (v) = M (v) N (v)T where M (v) ∑ (v) N (v)T is the singular value decomposition equivalent form of X (v) S (v)T G (v)T P T S (v) = W (v) + I, is the pre-constructed similarity graph matrix;

[0015] Solve the optimization problem of P: Obtain the optimal solution of the variable P as: where Let μ>0 be the positive penalty parameter, C be the Lagrange multiplier, Q be the auxiliary variable and P = Q, and c denote the number of rows of matrix P;

[0016] Solve the optimization problem of Q: The optimal solution of variable Q is obtained as: Q = (μP + C)(11 T + μI) -1 ;

[0017] Solve the optimization problem of α: The optimal solution of variable α is obtained as:

[0018] Among them,

[0019] The update formulas of C and μ are: Among them, ρ and μ 0 are constants;

[0020] Clustering process: Use the optimized optimal shared consistent representation matrix P to obtain the clustering results of the data, specifically including: According to If the j-th element value in the i-th column of P :,i is the largest, then the i-th sample is assigned to the j-th category. The clustering results of all samples can be obtained by finding the positions corresponding to the maximum element values in each column of the representation matrix P.

[0021] Furthermore, is a binary matrix of 0 and 1, used to retain the sample representations in matrix P corresponding to X (v) Matrix G (v) is constructed according to the view missing prior position index matrix Z. The specific construction method is as follows:

[0022]

[0023] Furthermore, preprocess the incomplete multi-view data with the given view missing prior position index matrix Z The specific steps include:

[0024] Missing view deletion: According to the view missing prior position index matrix Z, delete the missing examples in each view to obtain the set of non-missing data

[0025] Data normalization: Normalize The calculation method is Among them represents the i-th column vector of matrix X (v) ;

[0026] Local nearest neighbor graph Construction: For the non-missing data X of each perspective (v) , calculate the distance between each sample and its k nearest neighbor samples using the Gaussian kernel. The calculation method is where is one of the k nearest neighbors of the sample , and the other non-nearest neighbor elements in W (v) are set to 0;

[0027] Construct a transformation matrix according to the view missing prior position index matrix Z

[0028] In the second aspect of the present invention, an incomplete multi-view clustering system based on local structure and balance perception is provided. The system includes:

[0029] A model establishment unit, which is used for the clustering task of incomplete multi-view data, and designs an incomplete multi-view Figure 1 consistent clustering representation learning model with probability characteristics based on local structure and balance perception, specifically:

[0030]

[0031]

[0032] where represents the basis matrix of the v-th perspective, m v represents the feature dimension of the v-th view, d represents the dimension of the consistent representation space, P ∈ R d×n represents the shared consistent representation matrix of incomplete multi-view data, n represents the total number of samples of incomplete multi-view data, α = [α 1 ,..., α l is a learnable weight vector, 1 ∈ R d represents a d-dimensional column vector with all element values being 1, λ is a penalty term parameter, l represents the number of views, n v represents the number of non-missing samples of the v-th view, I is the identity matrix, I i,j represents the element value at the (i, j) row and column position of the identity matrix, represents the similarity relationship between the i-th sample and the j-th sample in the v-th view, represents the matrix X (v) 's i-th column vector, represents the matrix set composed of non-missing examples of the v-th view, represents the matrix G (v) 's j-th column vector, is a binary matrix of 0 and 1;

[0033] A data preprocessing unit, which is used for incomplete multi-view data with a given view missing prior position index matrix Z Perform preprocessing;

[0034] An optimization model unit for obtaining an incomplete multi-view consistent clustering representation learning model according to the preprocessed data Figure 1 Regarding the variables P, α and the introduced auxiliary variables Q, Lagrange multiplier C, and positive penalty parameter μ, design an alternating iterative optimization-based method to solve the variables to achieve the purpose of model optimization, where:

[0035] Solve for U (v) The optimization problem of: Obtain the variable U (v) The optimal solution of is U (v) = M (v) N (v)T where M (v) ∑ (v) N (v)T is X (v) S (v)T G (v)T P T The singular value decomposition of, S (v) = W (v) + I, is a pre-constructed similarity graph matrix;

[0036] Solve the optimization problem of P: The optimal solution of the variable P is: where μ > 0 is the positive penalty parameter, C is the Lagrange multiplier, Q is the auxiliary variable and P = Q, c represents the number of rows of the matrix P;

[0037] Solve the optimization problem of Q: The optimal solution of the variable Q is: Q = (μP + C)(11 T + μI) -1 ;

[0038] Solve the optimization problem of α: The optimal solution of the variable α is:

[0039] where,

[0040] The update formulas for C and μ are: where ρ and μ 0 are constants;

[0041] A clustering unit for obtaining the clustering result of the data by using the optimized optimal shared consistent representation matrix P, specifically including: According to If the i-th column of P :,iIf the value of the j-th element in [matrix] is the largest, then the i-th sample is assigned to the j-th category. By finding the position corresponding to the largest element value in each column of the representation matrix P, the clustering results of all samples can be obtained.

[0042] Furthermore, is a binary matrix of 0 and 1, used to retain the sample representations in matrix P corresponding to X (v) The matrix G (v) is constructed according to the view missing prior position index matrix Z. The specific construction method is as follows:

[0043]

[0044] Furthermore, the specific steps of the data preprocessing unit include:

[0045] Missing view deletion: According to the view missing prior position index matrix Z, delete the missing examples in each view to obtain the set of non-missing data

[0046] Data normalization: Perform normalization preprocessing on The calculation method is where represents the i-th column vector of matrix X (v) ;

[0047] Local nearest neighbor graph Construction: For the non-missing data X (v) of each view, use the Gaussian kernel to calculate the distance between each sample and its k nearest neighbor samples. The calculation method is where is one of the k nearest neighbors of the example , and the other non-nearest neighbor elements in W (v) are set to 0;

[0048] Construct the transformation matrix according to the view missing prior position index matrix Z

[0049] In the third aspect of the present invention, an incomplete multi-view clustering system based on local structure and balance perception is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned incomplete multi-view clustering method based on local structure and balance perception.

[0050] In the fourth aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, and when the instructions are executed by a processor, the processor is enabled to execute the above-mentioned incomplete multi-view clustering method based on local structure and balance perception.

[0051] An incomplete multi-view clustering method and system based on local structure and balance awareness provided by the present invention. The Local strUcture-andBalance-Aware Efficient Incomplete Multi-View Clustering (LUBA_EIMVC) aims at the efficient learning problem of incomplete multi-views, and designs an incomplete multi-view Figure 1 consistent clustering representation learning model with probability characteristics. This model obtains a unique clustering result by learning the consistent representation with probability characteristics between views. Each element in the consistent probability representation vector can directly reflect the probability that the corresponding sample belongs to a certain category. In addition, this model integrates geometric structure preservation and consistent representation learning into a very concise model, without introducing any additional constraint terms and penalty term parameters due to the introduction of geometric structure preservation characteristics, which not only makes the model more concise, but also reduces the burden of parameter tuning. In addition, in order to avoid samples being over-divided into a few categories, a balance awareness learning technique is introduced. The method of the present invention not only has the best and most stable clustering performance, but also is more computationally efficient than the current relatively advanced incomplete multi-view clustering methods. Specifically, the beneficial effects of the present invention include:

[0052] LUBA_EIMVC designs a novel balance-aware graph-regularized incomplete multi-view orthogonal matrix factorization model, which can not only exploit the local structure information of views to guide the optimization of the model, but also make full use of the available view information to learn the clustering consistency representation with probability characteristics;

[0053] Different from the existing method of using k-means to obtain clustering results, LUBA_EIMVC directly obtains a unique positive probability matrix shared by all views. In this matrix, each element can be regarded as the probability that a sample belongs to a certain category, thus solving the inaccurate clustering problem caused by k-means;

[0054] In order to avoid the problem that samples are overly concentrated in some classes or even one class during the process of the model optimizing the clustering results, a balance awareness constraint of the probability matrix is introduced, and a consistent representation matrix with clustering friendliness and probability characteristics is jointly learned. The clustering result of incomplete multi-view data can be directly obtained based on this matrix;

[0055] Due to the learning of the consistent representation with probability characteristics, the model designed by LUBA_EIMVC is an incomplete multi-view clustering model with interpretability, high efficiency and stable clustering results. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the incomplete multi-view clustering method based on local structure and balance perception in the embodiment of the present invention;

[0057] Figure 2 is the incomplete multi-view in the embodiment of the present invention Figure 1 It is a schematic diagram of the optimization learning process of the clustering representation learning model for the incomplete multi-view

[0058] Figure 3 It is a schematic diagram of the structure of the incomplete multi-view clustering system based on local structure and balance perception in the embodiment of the present invention;

[0059] Figure 4 It is an architecture diagram of a computer device in the embodiment of the present invention. Detailed implementation manners

[0060] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0061] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0062] The embodiment of the present invention provides the following embodiments for an incomplete multi-view clustering method and system based on local structure and balance perception:

[0063] Based on Embodiment 1 of the present invention

[0064] A schematic diagram of the incomplete multi-view clustering method based on local structure and balance perception in Embodiment 1 of the present invention is as Figure 1 shown, which is a schematic diagram of the Local Structure-and Balance-Aware Efficient Incomplete Multi-View Clustering (LUBA_EIMVC) method. The purpose of this method is to obtain a positive probability consistent representation matrix with an equilibrium clustering result from incomplete multi-view data, so as to output an interpretable and unique clustering result. The specific steps are as follows:

[0065] Model establishment: For the clustering task based on incomplete multi-view data, design an incomplete multi-view clustering representation learning model with probability characteristics based on local structure and balance perception, specifically: Figure 1 Specifically:

[0066]

[0067]

[0068] Among them, represents the basis matrix of the v-th view, m v represents the feature dimension of the v-th view, d represents the dimension of the consistent representation space, P ∈ R d×n represents the shared consistent representation matrix of the incomplete multi-view data, n represents the total number of samples of the incomplete multi-view data; α = [α 1 ,..., α l is a learnable weight vector, 1 ∈ R d represents a d-dimensional column vector with all element values being 1, α v represents the v-th element in the vector α, r is a positive integer not less than 2, represents the r-th power of the element α v in the vector α; λ is the penalty term parameter, l represents the number of views, n v represents the number of non-missing samples in the v-th view, I is the identity matrix, I i,j represents the element value at the (i, j)-th row and column position of the identity matrix, represents the similarity relationship between the i-th sample and the v-th view of the j-th sample, represents the i-th column vector of the matrix X (v) ; represents the matrix set composed of non-missing examples in the v-th view, represents the j-th column vector of the matrix G (v) ; is a binary matrix of 0 and 1;

[0069] In the specific implementation process, for the given incomplete multi-view data and the view missing position information matrix Z ∈ R l×n , Y (v) represents the matrix set composed of n examples of the v-th view, and its column vector represents the v-th view feature of the j-th sample. If the v-th view of the j-th sample is missing, the corresponding element Z v,j in the view missing position information matrix Z is 0; otherwise Z v,j = 1 indicates that the view of the corresponding sample is not missing. In the original data Among them, all the eigen - elements of the eigen - vector corresponding to the missing view can be identified by "NaN". For the clustering task of such incomplete multi - view data, the embodiment designs a locally - structured and balance - aware incomplete multi - view Figure 1 probabilistic clustering representation learning model. In formula (1), matrix X (v) can be directly obtained by deleting the columns corresponding to the vectors represented as "NaN" in the original data Y (v) . In formula (1), represents the basis matrix of the v - th perspective, and the value of this variable can be obtained by optimizing model (1), where d represents the dimension of the consistent representation space, usually set to the number of categories into which the data is expected to be partitioned. 1≥P≥0 represents that the range of all element values in matrix P is [0,1]. is a pre - constructed similarity graph matrix, and its element represents the similarity relationship between the i - th sample and the j - th sample in the v - th view. Its specific construction method is as follows: 1) Calculate the Gaussian distance between non - missing examples within the v - th view, and the calculation method is: 2) For the i - th non - missing example, sort the distances between it and the other n v - 1 examples, and only retain the Gaussian distances corresponding to the first k smallest examples for each example in W (v) , and set other elements to 0.

[0070] In the preferred embodiment, in formula (1), is a binary matrix of 0 and 1, used to retain the sample representations in matrix P corresponding to X (v) . Matrix G (v) is constructed according to the view - missing prior position index matrix Z, and the specific construction method is as follows:

[0071]

[0072] In model (1), the orthogonality constraint U (v) U (v)T U (v) =I on the basis matrix U can avoid the problem of clustering center degradation. The constraint term is a balance - aware constraint, which can avoid the problem of over - clustering into a small number of categories, that is, classifying the data with c categories into categories. Compared with the binary clustering labels obtained by kmeans, the constraint P T 1 = 1 introduced in model (1) results in a consistent probability matrix, which increases the degrees of freedom of basis matrix learning and consistent representation learning. In model (1), is the novel graph - embedding multi - view designed by the method of the present invention Figure 1Regarding the representation learning item, the innovation and distinctiveness from other methods are mainly reflected in that the present invention incorporates the graph embedding structure preservation property and the shared consistent representation learning of incomplete multi-views into a very concise model without any hyperparameters, and can obtain a structured discriminative consistent representation.

[0073] Data preprocessing: For the incomplete multi-view data with a missing prior position index matrix Z for a given view perform preprocessing;

[0074] In the preferred embodiment, for the incomplete multi-view data with a missing prior position index matrix Z for a given view perform preprocessing, and the specific steps include:

[0075] Missing view deletion: According to the view missing prior position index matrix Z, delete the missing examples in each view to obtain an un-missing data set

[0076] Data normalization: For perform normalization preprocessing, and the calculation method is where represents the i-th column vector of matrix X (v) ;

[0077] Local nearest neighbor graph Construction: For the un-missing data X of each view (v) , use the Gaussian kernel to calculate the distance between each sample and its k nearest neighbor samples, and the calculation method is where is one of the k nearest neighbors of the example , and other non-nearest neighbor elements in W (v) are set to 0;

[0078] Use Equation (2) to construct a transformation matrix according to the view missing prior position index matrix Z

[0079] Optimization model: According to the preprocessed data and the designed incomplete multi-view Figure 1 consistent clustering representation learning model (1), for the variables P, α and the introduced auxiliary variables Q, Lagrange multipliers C and positive penalty parameter μ contained in the model, design an alternating iterative optimization-based method to solve the variables to achieve the purpose of model optimization.

[0080] In the specific implementation process, model (1) contains multiple variables such as P and α, and design an alternating iterative optimization-based method to solve these variables. First, let S (v) = W (v) + I and introduce an auxiliary variable Q and let P = Q, as follows:

[0081]

[0082] The augmented Lagrangian function of problem (3) can be expressed as:

[0083]

[0084] where μ > 0 is the positive penalty parameter; C is the Lagrange multiplier. denotes the "Frobenius" norm of matrix A∈Rm×n and its calculation method is where A i,j is the (i, j)-th element of matrix A.

[0085] Then, by iteratively solving and optimizing the following five problems one by one, the optimal solutions of these variables can be obtained:

[0086] Step 1: Solve for U (v) , when solving for U (v) , the other variables to be solved can be considered as known, and then the following optimization sub-problem for variable U (v) can be obtained:

[0087]

[0088] According to the constraint U (v)T U (v) = I, problem (5) can be simplified to:

[0089]

[0090] where D (v) is a diagonal matrix. Since matrix S (v) has a symmetric structure, the calculation formula for D (v) is

[0091] According to equation (6), the following optimization problem equivalent to problem (5) can be obtained:

[0092]

[0093] Let the singular value decomposition (SVD) of X (v) S (v)T G (v)T P T be M (v) ∑ (v) N (v)T , then the optimal solution of problem (7) is U (v) = M (v) N (v)T, the singular value decomposition operation can be directly obtained by calling the "svd" function in the Matlab software.

[0094] Step 2: Solve for P. Considering the variables other than P as known quantities, the following optimization sub-problem for the variable P can be obtained:

[0095]

[0096] Problem (8) can be simplified to:

[0097]

[0098] In the formula It can be found that matrix A is a diagonal matrix with all diagonal elements being positive. According to Problem (9) can be transformed into the following equivalent form:

[0099]

[0100] Problem (10) can be regarded as n independent optimization problems. Therefore, the optimal solution of the variable P can be obtained by optimizing the problem with respect to P :,i column by column, as shown below:

[0101]

[0102] In the formula, the function max(a, 0) means setting the element a less than 0 to 0. c represents the number of rows of matrix P.

[0103] Step 3: Solve for Q. After fixing the variables other than Q, the optimization problem degenerates to:

[0104]

[0105] By calculating the partial derivative of problem (12) with respect to the variable Q and then setting it to 0, we can obtain:

[0106] Q = (μP + C)(11 T + μI) -1 (13)

[0107] Step 4: Solve for α. Let and fix the variables independent of the variable α, the following optimization problem for the variable α can be obtained:

[0108]

[0109] Solving problem (14), the optimal solution of the variable α can be obtained as:

[0110]

[0111] where r is a positive integer not less than 2.

[0112] Step 5: Update C and μ. The update formulas for C and μ are as follows:

[0113]

[0114] In the formula, ρ and μ 0 are constants, where μ 0 is usually set to a relatively large value such as 10 8 , and ρ is usually set to a value greater than 1.

[0115] A specific example of the model optimization process is shown in Algorithm 1 below. For LUBA_EIMVC, the label of the i-th sample can be obtained directly through direct acquisition.

[0116]

[0117]

[0118] The complete optimization flowchart for the model (1) is as shown in Figure 2 , where the initialization step mainly includes in Algorithm 1: Initialize as a random arbitrary orthogonal matrix. Initialize α = 1, μ 0 = 10 8 , μ = 0.01, ρ = 1.08. Figure 2 The convergence criterion in t is |loss t-1 - loss -5 | < 10 t , where loss t-1 and loss

[0119]

[0120] represent the objective loss values at the t-th step and the (t - 1)-th step respectively, and the calculation formula is: Clustering process: Using the optimized optimal shared consistent representation matrix P, according to :,i If the value of the j-th element in the i-th column of P is the largest, then the i-th sample is assigned to the j-th category. By obtaining the position corresponding to the largest element value in each column of the representation matrix P, the clustering results of all samples can be obtained. d represents the number of rows of the matrix P, and generally can be set to the number of clustering categories c.

[0121] Because the elements in the matrix P of the present invention represent the probability that each sample belongs to a certain class, the clustering results of the data can be directly obtained according to , that is, if the i-th column of P :,iIf the value of the j-th element in [matrix] is the largest, then the i-th sample is classified into the j-th category. By finding the position corresponding to the largest element value in each column of matrix P, the clustering results of all samples can be obtained.

[0122] The method of the present invention designs a new model for clustering incomplete multi-view data with fast speed, stability and interpretability for the problem of clustering incomplete multi-view data with missing views in application scenarios of all walks of life. The unique features compared with previous models are as follows: The model is concise and distinctive: A distinctive and concise "incomplete multi-view consistent representation learning term" is proposed, which integrates local structure embedding and incomplete multi-view consistent representation learning into an optimization term; The model is interpretable: Each constraint term of the model has its meaning and value, and each element value of the obtained shared consistent representation is a representation of the clustering result. Therefore, both the model itself and the model output results are interpretable; The model has a unique stable solution: Different from traditional methods that require additional kmeans clustering to obtain unstable clustering results, the method of the present invention can directly obtain the unique clustering result of the data according to the unique output "consistent representation probability matrix P" of the model. The method of the present invention can not only obtain higher clustering accuracy, but also has the least time overhead and the highest efficiency. Figure 1 consistent representation learning term", which embeds local structure and incomplete multi-view Figure 1 consistent representation learning into an optimization term; The model is interpretable: Each constraint term of the model has its meaning and value, and each element value of the obtained shared consistent representation is a representation of the clustering result. Therefore, both the model itself and the model output results are interpretable; The model has a unique stable solution: Different from traditional methods that require additional kmeans clustering to obtain unstable clustering results, the method of the present invention can directly obtain the unique clustering result of the data according to the unique output "consistent representation probability matrix P" of the model. The method of the present invention can not only obtain higher clustering accuracy, but also has the least time overhead and the highest efficiency.

[0123] Based on Embodiment 2 of the present invention

[0124] An incomplete multi-view clustering system 300 based on local structure and balance perception provided by Embodiment 2 of the present invention can execute the incomplete multi-view clustering method based on local structure and balance perception provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. The system can be implemented in the form of software and / or hardware (integrated circuit), and is generally integrated in a server or a terminal device.

[0125] Figure 3 is a schematic structural diagram of an incomplete multi-view clustering system 300 based on local structure and balance perception in Embodiment 2 of the present invention. Referring to Figure 3 , the incomplete multi-view clustering system 300 based on local structure and balance perception in the embodiment of the present invention specifically may include:

[0126] A model building unit 310, configured to design an incomplete multi-view consistent clustering representation learning model with probability characteristics based on local structure and balance perception for the clustering task of incomplete multi-view data, specifically: Figure 1 consistent clustering representation learning model, specifically:

[0127]

[0128]

[0129] Among them, represents the basis matrix of the v-th perspective, m v represents the feature dimension of the v-th view, d represents the dimension of the consistent representation space, P ∈ R d×n represents the shared consistent representation matrix of the incomplete multi-view data, n represents the total number of samples of the incomplete multi-view data, α = [α 1 ,..., α l is a learnable weight vector, 1 ∈ R d represents a d-dimensional column vector with all element values being 1, λ is the penalty term parameter, l represents the number of views, n v represents the number of non-missing samples in the v-th view, I is the identity matrix, I i,j represents the element value at the (i, j)-th row and column position of the identity matrix, represents the similarity relationship between the i-th sample and the v-th view of the j-th sample, represents matrix X (v) 's i-th column vector, represents the matrix set composed of non-missing examples in the v-th view, represents matrix G (v) 's j-th column vector, is a binary matrix of 0 and 1; r is a positive integer not less than 2.

[0130] The data preprocessing unit 320 is used to preprocess the incomplete multi-view data of the given view missing prior position index matrix Z for preprocessing;

[0131] The optimization model unit 330 is used to design an alternating iteration optimization-based method to solve the variables for the variables Figure 1 P, α and the introduced auxiliary variables Q, Lagrange multipliers C and positive penalty parameter μ according to the preprocessed data and the designed incomplete multi-view consistent clustering representation learning model, so as to achieve the purpose of model optimization, where:

[0132] Solve the optimization problem of U (v) : Get the optimal solution of the variable U (v) as U (v) = M (v) N (v)T , where M (v) ∑ (v) N (v)T is the equivalent form of the singular value decomposition of X (v) S (v)T G (v)T P T , S (v) = W (v)+I, is a pre - constructed similarity graph matrix;

[0133] Solve the optimization problem of P: The optimal solution of variable P is: where μ > 0 is a positive penalty parameter, C is a Lagrange multiplier, Q is an auxiliary variable and P = Q;

[0134] Solve the optimization problem of Q: The optimal solution of variable Q is: Q=(μP + C)(11 T + μI) -1 ;

[0135] Solve the optimization problem of α: The optimal solution of variable α is:

[0136] where,

[0137] The update formulas of C and μ are: where ρ and μ 0 are constants;

[0138] The clustering unit 340 is used to obtain the clustering result of the data by using the optimized optimal shared consistent representation matrix P. According to If the j - th element value in the i - th column of P :,i is the largest, then the i - th sample is classified into the j - th category. The clustering result of all samples can be obtained by finding the position corresponding to the largest element value in each column of the representation matrix P.

[0139] Furthermore, is a binary matrix of 0 and 1, which is used to retain the sample representation in matrix P corresponding to X (v) Matrix G (v) is constructed according to the view - missing prior position index matrix Z. The specific construction method is as follows:

[0140]

[0141] Furthermore, the specific steps of the data pre - processing unit 320 include:

[0142] Missing view deletion: According to the view - missing prior position index matrix Z, delete the missing examples in each view to obtain the non - missing data set

[0143] Data normalization: Perform normalization pre - processing on The calculation method is where represents matrix X(v) the i-th column vector of

[0144] local nearest neighbor graph Construction: For the non-missing data X of each view (v) , use the Gaussian kernel to calculate the distance between each sample and its k nearest neighbor samples, and the calculation method is where is a sample one of the k nearest neighbors of, and the other non-nearest neighbor elements in W (v) are set to 0;

[0145] Construct the transformation matrix according to the view missing prior position index matrix Z

[0146] In addition to the above four units, the system 300 may further include other components. However, since these components are not related to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.

[0147] The specific working process of an incomplete multi-view clustering system 300 based on local structure and balance awareness refers to the description of Embodiment 1 of the incomplete multi-view clustering method based on local structure and balance awareness above, and will not be repeated here.

[0148] Based on Embodiment 3 of the present invention

[0149] The system according to an embodiment of the present invention may also be implemented with the aid of Figure 4 the architecture of the computing device shown. Figure 4 shows the architecture of the computing device. As Figure 4 shown, a computer system 410, a system bus 430, one or more CPUs 440, an input / output component 420, a memory 450, etc. The memory 450 may store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the method of Embodiment 1. Figure 4 The architecture shown is only exemplary, and when implementing different devices, one or more components in Figure 4 are adjusted according to actual needs.

[0150] Based on Embodiment 4 of the present invention

[0151] Embodiments of the present invention may also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium according to Embodiment 4. When the computer-readable instructions are run by a processor, the incomplete multi-view clustering method based on local structure and balance awareness according to Embodiment 1 of the present invention described with reference to the above drawings can be executed.

[0152] In the embodiments of the present invention, for the above-mentioned incomplete multi-view clustering method, system and storage medium based on local structure and balance awareness, Examples 1-4 are used to train and test the proposed incomplete multi-view clustering method and system based on local structure and balance awareness. Table 2 shows the average clustering accuracy obtained on the BBCSport, Caltech101 and 3Sources data sets with a view missing rate of 30%. Among them, MIC and DAIMC are the current incomplete multi-view clustering methods for manifolds, and the results are shown in Table 2:

[0153] Table 2

[0154] Dataset MIC DAIMC The present invention BBCSport 46.21±4.71 63.45±1.97 78.79±3.02 Caltech101 20.12±0.75 25.15±0.31 27.63±0.90 3Sources 47.69±7.61 52.43±6.63 71.83±7.37

[0155] Table 3 shows the execution time (unit: seconds) on the BBCSport, Caltech101 and 3Sources data sets with a view missing rate of 30%. Among them, MIC and DAIMC are the current incomplete multi-view clustering methods for manifolds, and the results are shown in Table 3:

[0156] Table 3

[0157] Dataset MIC DAIMC The present invention BBCSport 3.843 148.501 2.183 Caltech101 <![CDATA[1.407×10 4 > <![CDATA[1.861×10 3 > 129.541 3Sources 5.912 563.780 4.967

[0158] Using Examples 1-4 and the above performance analysis, the present invention proposes an incomplete multi-view clustering method and system based on local structure and balance awareness. The incomplete multi-view clustering based on local structure and balance awareness (LUBA_EIMVC, Local strUcture-and Balance-Aware Efficient Incomplete Multi-ViewClustering) aims at the efficient learning problem of incomplete multi-views and designs an incomplete multi-view with probabilistic characteristics Figure 1To a clustering representation learning model, which obtains a unique clustering result by learning consistent representations with probabilistic characteristics among views, where each element in the consistent probability representation vector can directly reflect the probability that the corresponding sample belongs to a certain category. Additionally, the model integrates geometric structure preservation and consistent representation learning into a very concise model, without introducing any additional constraint terms and penalty term parameters due to the introduction of geometric structure preservation characteristics, not only making the model more concise but also reducing the parameter tuning burden. Furthermore, in order to avoid samples being overly divided into a few categories, a balance-aware learning technique is introduced. The method of the present invention not only has the best and most stable clustering performance, but also is more computationally efficient than the current relatively advanced incomplete multi-view clustering methods. Specifically, the beneficial effects of the present invention include: LUBA_EIMVC designs a novel balance-aware graph-regularized incomplete multi-view orthogonal matrix factorization model, which can not only exploit the local structure information of views to guide the optimization of the model, but also make full use of the available view information to learn clustering consistency representations with probabilistic characteristics; different from the existing methods that use k-means to obtain clustering results, LUBA_EIMVC directly obtains a unique positive probability matrix shared by all views, in which each element can be regarded as the probability that the sample belongs to a certain category, thus solving the problem of inaccurate clustering results caused by k-means; in order to avoid the problem that samples are overly concentrated in some categories or even a single category during the process of the model optimizing the clustering results, a balance-aware constraint of the probability matrix is introduced to jointly learn a consistent representation matrix with clustering friendliness and probabilistic characteristics, and the clustering result of incomplete multi-view data can be directly obtained based on this matrix; due to the learning of consistent representations with probabilistic characteristics, the model designed by LUBA_EIMVC is an incomplete multi-view clustering model with interpretability, high efficiency, and stable clustering results.

[0159] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An incomplete multi-view clustering method based on local structure and balance awareness, Characterized in that, It includes the following steps: Model establishment: For the clustering task of incomplete multi-view data, design an incomplete multi-view consistent clustering representation learning model with probability characteristics based on local structure and balance awareness. The specific model is: Among them, represents the basis matrix of the v-th perspective, m v represents the feature dimension of the v-th view, d represents the dimension of the consistent representation space, P ∈ R d×n represents the shared consistent representation matrix of the incomplete multi-view data, n represents the total number of samples of the incomplete multi-view data, α = [α 1 ,..., α l is a learnable weight vector, 1 ∈ R d represents a d-dimensional column vector with all element values being 1, α v represents the v-th element in the vector α, r is a positive integer not less than 2, represents the r-th power of the element α v in the vector α, λ is the penalty term parameter, l represents the number of views, n v represents the number of non-missing samples in the v-th view, I is the identity matrix, I i,j represents the element value at the (i, j)-th row and column position of the identity matrix, represents the similarity relationship between the i-th sample and the v-th view of the j-th sample, represents the i-th column vector of the matrix X (v) , represents the matrix set composed of non-missing examples in the v-th view, represents the j-th column vector of the matrix G (v) , is a binary matrix of 0 and 1; Data preprocessing: Preprocess the incomplete multi-view data with the prior position index matrix Z missing in the given view for preprocessing; Optimized model: Based on the preprocessed data and the incomplete multi-view consistent clustering representation learning model, for the variables P, α, and the introduced auxiliary variables Q, Lagrange multipliers C, and positive penalty parameter μ, design an alternating iterative optimization-based method to solve the variables to achieve the purpose of model optimization, where: Solve for U (v) The optimization problem of: Obtain the variable U (v) The optimal solution of is U (v) = M (v) N (v)T where M (v) ∑ (v) N (v)T is the equivalent form of the singular value decomposition of X (v) S (v)T G (v)T P T and S (v) = W (v) + I, is the pre-constructed similarity graph matrix; Solve the optimization problem of P: The optimal solution of the variable P is obtained as: where μ > 0 is a positive penalty parameter, C is a Lagrange multiplier, Q is an auxiliary variable and P = Q, c represents the number of rows of the matrix P; Solve the optimization problem of Q: The optimal solution of the variable Q is obtained as: Q = (μP + C)(11 T + μI) -1 ; Optimization problem for solving α: The optimal solution of the variable α is obtained as: Among them, The update formulas for C and μ are as follows: where ρ and μ 0 are constants; Clustering process: The clustering results of the data are obtained by using the optimal shared consistent representation matrix P obtained after optimization, which specifically includes: According to If the value of the j-th element in the i-th column of P :,i is the largest, then the i-th sample is assigned to the j-th category. The clustering results of all samples can be obtained by finding the positions corresponding to the maximum element values in each column of the representation matrix P; The multi-view data is any one of the BBCSport, Caltech101, and 3Sources datasets.

2. The incomplete multi-view clustering method based on local structure and balance awareness according to claim 1, Characterized in that, is a binary matrix of 0 and 1, used to retain the sample representation corresponding to X in matrix P (v) matrix G (v) is constructed according to the view missing prior position index matrix Z, and the specific construction method is as follows:

3. The incomplete multi-view clustering method based on local structure and balance awareness according to claim 2, Characterized in that, Incomplete multi-view data lacking the prior position index matrix Z for a given view Perform preprocessing, and the specific steps include: Missing view deletion: According to the prior position index matrix Z of missing views, missing examples in each view are deleted to obtain a set of non-missing data Data normalization: For perform preprocessing of normalization, and the calculation method is where represents the i-th column vector of matrix X (v) ; Local nearest neighbor graph Construction: For the non-missing data X of each view (v) , use the Gaussian kernel to calculate the distance between each sample and its k nearest neighbor samples. The calculation method is where is one of the k nearest neighbors of the sample , and the other non-nearest neighbor elements in W (v) are set to 0; Construct a transformation matrix according to the prior position index matrix Z missing from the view 4. An incomplete multi-view clustering system based on local structure and balance awareness, Characterized in that, The system includes: A model establishment unit for designing an incomplete multi-view consistent clustering representation learning model with probability characteristics based on local structure and balance awareness for the clustering task of incomplete multi-view data. The specific model is: Among them, represents the basis matrix of the v-th perspective, m v represents the feature dimension of the v-th view, d represents the dimension of the consistent representation space, P ∈ R d×n represents the shared consistent representation matrix of the incomplete multi-view data, n represents the total number of samples of the incomplete multi-view data, α = [α 1 ,..., α l is a learnable weight vector, 1 ∈ R d represents a d-dimensional column vector with all element values being 1, α v represents the v-th element in the vector α, r is a positive integer not less than 2, represents the r-th power of the element α v in the vector α, λ is the penalty term parameter, l represents the number of views, n v represents the number of non-missing samples in the v-th view, I is the identity matrix, I i,j represents the element value at the (i, j)-th row and column position of the identity matrix, represents the similarity relationship between the i-th sample and the v-th view of the j-th sample, represents the i-th column vector of the matrix X (v) , represents the matrix set composed of non-missing examples in the v-th view, represents the j-th column vector of the matrix G (v) , is a binary matrix of 0 and 1; A data preprocessing unit for preprocessing incomplete multi-view data with a missing prior position index matrix Z for a given view for preprocessing; Optimization model unit, which is used to design an alternating iteration optimization-based method to solve variables according to the preprocessed data and the incomplete multi-view consistent clustering representation learning model, aiming at the variables contained in the model For variables P, α, and the introduced auxiliary variables Q, Lagrange multipliers C, and positive penalty parameter μ, the purpose of model optimization is achieved by solving the variables Solve for U (v) The optimization problem of: Obtain the variable U (v) The optimal solution of is U (v) = M (v) N (v)T where M (v) ∑ (v) N (v)T is the equivalent form of the singular value decomposition of X (v) S (v)T G (v)T P T and S (v) = W (v) + I, is a pre-constructed similarity graph matrix; Solve the optimization problem of P: The optimal solution of the variable P is obtained as: where μ > 0 is a positive penalty parameter, C is a Lagrange multiplier, Q is an auxiliary variable and P = Q, c represents the number of rows of the matrix P; Solve the optimization problem of Q: The optimal solution of the variable Q is obtained as: Q = (μP + C)(11 T + μI) -1 ; Optimization problem for solving α: The optimal solution for variable α is: Among them, The update formulas for C and μ are as follows: where ρ and μ 0 are constants; A clustering unit, which is used to obtain the clustering result of data by using the optimal shared consistent representation matrix P obtained after optimization, specifically including: According to If the j-th element value in the i-th column of P :,i is the largest, then the i-th sample is classified into the j-th category. By obtaining the positions corresponding to the maximum element values in each column of the representation matrix P, the clustering results of all samples can be obtained; The multi-view data is any one of the BBCSport, Caltech101, and 3Sources datasets.

5. The incomplete multi-view clustering system based on local structure and balance awareness according to claim 4, Characterized in that, is a binary matrix of 0 and 1, used to retain the sample representation corresponding to X in matrix P (v) matrix G (v) is constructed according to the view missing prior position index matrix Z, and the specific construction method is as follows:

6. The incomplete multi-view clustering system based on local structure and balance awareness according to claim 5, Characterized in that, The specific steps of the data preprocessing unit include: Missing view deletion: According to the prior position index matrix Z of missing views, missing examples in each view are deleted to obtain a set of non-missing data Data normalization: For perform preprocessing of normalization, and the calculation method is where represents the i-th column vector of matrix X (v) ; Local nearest neighbor graph Construction: For the non-missing data X of each view (v) , use the Gaussian kernel to calculate the distance between each sample and its k nearest neighbor samples, and the calculation method is where is a sample and one of the k nearest neighbors of, W (v) Set other non-nearest neighbor elements in to 0; Construct a transformation matrix according to the prior position index matrix Z missing from the view 7. An incomplete multi-view clustering system based on local structure and balance awareness, Characterized in that, It includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the incomplete multi-view clustering method based on local structure and balance awareness according to any one of claims 1 to 3.

8. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the incomplete multi-view clustering method based on local structure and balance awareness according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Multi-view subspace clustering method for self-weighted fusion of local and global information

    CN113554082A

  • Incomplete multi-view clustering method based on missing graph reconstruction and adaptive neighbor

    CN113947135A