Fine-grained multi-view image optimization method based on augmented lagrangian method
Patent Information
- Application Number
- CN202311507292.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-11-10
AI Technical Summary
多视图锚点图聚类可能非常困难,因为锚点在特征维度上不一致
本发明不是通过融合多个相似图来获得共识图,而是开发了一种细粒度的融合策略,该策略在样本水平上考虑融合过程。
Smart Images

Figure CN117593552B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-view subspace clustering technology, and specifically to a fine-grained multi-view image optimization method based on the augmented Lagrange method. Background Technology
[0002] The proliferation of multi-view data sources has increased the demand for methods capable of extracting valuable insights from numerous advantageous locations. For example, autonomous driving requires both RGB video and corresponding dense 3D point clouds to construct more accurate real-time traffic information. Disease diagnosis needs to consider health history data, physical examinations, and blood tests. Meanwhile, image datasets can be described using multiple different descriptors, such as LBP, SIFT, and HOG. However, traditional single-view clustering methods excel at separating data in isolated patterns and often fail to encapsulate comprehensive knowledge embedded across multiple perspectives.
[0003] Among existing multi-view clustering (MVC) techniques, graph-based clustering methods have attracted considerable attention due to their efficiency and ease of implementation. Typically, graph-based MVC methods first learn a similarity graph S for each view. (v) Then, different fusion strategies are deployed on these view-oriented graphs to obtain an optimal consensus graph S, which can be further processed by two mainstream techniques. One is to restrict the rank of S to n - k, where n is the number of data points and k is the number of clusters. Since there are now k connection components about k clusters in the optimal consensus graph S, no post-processing is needed to obtain the clustering results. The second is to perform spectral clustering or... k -means, to get the final result.
[0004] Utilizing the diffusion process, Tang et al. proposed a parameter-free MVC method based on transverse graph diffusion in 2020. Since the original graph may be noisy or incomplete and cannot be directly applied, Pan et al. proposed a multi-view comparative graph clustering method to learn the consensus graph in 2021. Multi-view anchor graph clustering can be very difficult because anchors are inconsistent in feature dimensions. To address this, Siwei et al. proposed a generalized and flexible anchor graph fusion framework called Fast Multi-View Anchor-Correspondence Clustering in 2022. To fully utilize multi-view information, Wang et al. designed a specific graph learning method in 2022 by introducing graph regularization and local structure fusion patterns. Li et al. proposed a high-order correlation-preserving incomplete multi-view subspace clustering method in 2022, effectively recovering the subspace structure of missing views and incomplete multi-view data. To simultaneously consider inter-view similarity and intra-view similarity, Xia et al. proposed a variance-based bipartite construction decorrelation anchor selection strategy in 2023. To extract common information from high-level views and reduce the influence of non-homogeneous edges, Ling et al. proposed a bi-labeled guided graph improvement for multi-view graph clustering in 2023.
[0005] While the aforementioned work has demonstrated encouraging results, it is worth noting that existing graph-based multi-view clustering methods primarily focus on obtaining a consensus graph by fusing multiple view-oriented similar graphs in an inter-view manner. Almost none of them consider the intersection of heterogeneous information from a fine-grained perspective, let alone construct a self-consistent model that should not introduce any unexplained intermediate variables. Summary of the Invention
[0006] To address the aforementioned shortcomings in the prior art, this invention provides a fine-grained multi-view image optimization method based on the augmented Lagrange method.
[0007] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A fine-grained multi-view image optimization method based on the augmented Lagrange method includes the following steps: S1. Collect view sample data points and construct a multi-view data matrix based on each view sample data point; S2. Using the subspace correlation assumption, the subspace representation of the original data is obtained. Based on the augmented Lagrangian function method, the subspace representations of all the original data are combined into a third-order tensor and sliced along the view direction to obtain the topological representation of each sample in all views. S3. Calculate the optimal multi-view fusion graph and iteratively optimize the objective function to obtain the optimal multi-view subspace consensus representation; S4. Cluster the consensus representation of the optimal multi-view subspace using spectral clustering.
[0008] Furthermore, the Zeng Guang Lagrange function in S2 is expressed as:
[0009] In the formula, For the first A data matrix for each view, where n is the total number of sample data points for each view. For the first Data dimensions of each view; Operator for calculating the sum of squares of the elements of a matrix; Let S be the similarity matrix on the v-th view; S is the consensus similarity graph. Let i be the transpose of the i-th row of the consensus similarity graph; For Lagrange multipliers; As an auxiliary variable, It is the i-th slice of a third-order tensor formed by stacking auxiliary vectors; This is the weight matrix. For the first i The transpose of the weight vectors corresponding to each sample; This is the penalty term for the Lagrange function; They are respectively The hyperparameters of the constraint terms of S.
[0010] Furthermore, the topological representation of all views in S2 is as follows:
[0011] In the formula, Let v be the similarity matrix on the v-th view. For the first A data matrix for each view, where n is the total number of sample data points for each view. For the first Data dimensions of each view; It is a hyperparameter; This is a penalty item; For Lagrange multipliers; As an auxiliary variable; It is the identity matrix; superscript This is the matrix transpose symbol.
[0012] Furthermore, the process of S3 calculating the optimal multi-view fusion graph involves solving for each auxiliary variable separately. The weight matrix W and the consensus similarity graph S.
[0013] Furthermore, in S3, each auxiliary variable is solved. The specific method is as follows:
[0014] In the formula, It is the i-th slice of a third-order tensor formed by stacking auxiliary vectors; It is a hyperparameter; For the first i The weight vector corresponding to each sample for transpose; This is a penalty item; It is the identity matrix. Let i be the transpose of the i-th row of the consensus similarity graph; For the similarity matrix on all views The i-th slice in the spliced tensor; Let i be the i-th slice of the Lagrange multiplier.
[0015] Furthermore, the specific method for solving the weight matrix W in S3 is as follows:
[0016] In the formula, It is a process variable and , for The transpose of the matrix; It is an identity matrix.
[0017] Furthermore, the solution method for the consensus similarity graph S in S3 is as follows:
[0018] In the formula, The i-th row of the consensus similarity graph; λ is the hyperparameter of the sample-level fusion term; λ is the hyperparameter of the constraint term of the consensus graph S; For the first i The transpose of the weight vector corresponding to each sample It is the i-th slice of a third-order tensor formed by stacking auxiliary vectors.
[0019] Furthermore, the clustered representation of the optimal multi-view subspace consensus representation in S4 is as follows:
[0020] In the formula, Let v be the similarity matrix on the v-th view. For the similarity matrix on all views The i-th slice in the spliced tensor; S is an auxiliary variable; S is the consensus similarity graph. Let i be the transpose of the i-th row of the consensus similarity graph; This is the weight matrix. For the first i The transpose of the weight vectors corresponding to each sample; For the first A data matrix for each view, where n is the total number of sample data points for each view. For the first Data dimensions of each view; Operator for calculating the sum of squares of the elements of a matrix; Similarity matrix Hyperparameters of constraint terms; λ is the hyperparameter of the sample-level fusion term; λ is the hyperparameter of the constraint term of the consensus graph S. The present invention has the following beneficial effects: Instead of obtaining a consensus graph by fusing multiple similar graphs, this invention develops a fine-grained fusion strategy that considers the fusion process at the sample level.
[0021] The method of this invention is self-consistent, meaning that an efficient augmented Lagrangian method is introduced into the model of this invention so that it can obtain the optimal solution, unlike previous methods which involved an inexplicable intermediate variable. Therefore, self-consistency can be better maintained.
[0022] Extensive experiments and results on multiple benchmark datasets demonstrate the superiority of the method proposed in this invention. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the fine-grained multi-view image optimization method based on the augmented Lagrange method of the present invention.
[0024] Figure 2 This refers to the clustering results of the Yale dataset in this embodiment of the invention; Figure 3 This refers to the clustering results of the Citeseer dataset in this embodiment of the invention; Figure 4 This refers to the clustering results of the Cora dataset in this embodiment of the invention; Figure 5 This refers to the clustering results of the BBC dataset in this embodiment of the invention; Figure 6 This refers to the clustering results of the ORL dataset in this embodiment of the invention; Figure 7 This refers to the clustering results of the Cornell dataset in this embodiment of the invention; Figure 8This is a clustering NMI result diagram with different parameter combinations on the ORL dataset in an embodiment of the present invention, where af corresponds to the constraint term hyperparameters of the consensus graph S set to 700, 800, 900, 1000, 1100 and 1200 respectively. Figure 9 This is a graph showing the change of the target value with the number of iterations on the Cora and Cornell datasets in an embodiment of the present invention, where a represents the Cora dataset and b represents the Cornell dataset. Figure 10 The diagram below illustrates the visualization results of different datasets used in the embodiments of the present invention, where ac represents the visualization of the Yale dataset, HW dataset, and ORL dataset used, respectively. Detailed Implementation
[0025] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0026] A fine-grained multi-view image optimization method based on augmented Lagrange method, such as Figure 1 As shown, it includes the following steps: S1. Collect view sample data points and construct a multi-view data matrix based on each view sample data point; In this embodiment of the invention, collecting view sample data points can construct a sample dataset for each view, satisfying the following requirements: : And construct a multi-view dataset based on the sample datasets of each view. ,satisfy: Where n is the total number of sample data points for each view. For multi-view datasets The v One view, For the first v The data dimensions of each view.
[0027] S2. Using the subspace correlation assumption, the subspace representation of the original data is obtained. Based on the augmented Lagrangian function method, the subspace representations of all the original data are combined into a third-order tensor and sliced along the view direction to obtain the topological representation of each sample in all views. The fine-grained fusion strategy of this invention emphasizes the consistency of heterogeneous information, primarily by ensuring that local structures are roughly the same. Therefore, a weight is assigned to each sample in each view to more accurately calculate the sample-wise consensus graph, thereby suppressing the inconsistencies introduced by view-wise fusion. Thus, the fine-grained fusion strategy can be modeled as:
[0028] Where i is the sample index, n is the total number of samples, and S is the consensus similarity graph. For the i-th row of S, This is the weight matrix. For the first i The transpose of the weight vector corresponding to each sample. Meanwhile, It is composed of all { A tensor formed by concatenating v=1, …, m yes The i-th slice, These are two hyperparameters. Therefore, we can now obtain a consensus graph S that can describe more of the underlying data structures.
[0029] Inspired by the augmented Lagrange method, this invention first introduces auxiliary variables for each view. Transform it into the following equivalent problem:
[0030] for The i-th slice, For m The stacked third-order tensors allow the present invention to deploy an ALM to sort the aforementioned problems. The corresponding augmented Lagrangian function is:
[0031] in It is a penalty item. It is a Lagrange multiplier.
[0032] S3. Calculate the optimal multi-view fusion graph and iteratively optimize the objective function to obtain the optimal multi-view subspace consensus representation; This invention reduces the dimensionality of the feature space by employing subspace learning techniques and obtains a subspace representation matrix for each view. : For each The present invention requires solving:
[0033] Differentiating the above equation and setting it to zero, we have:
[0034] For each The present invention requires solving:
[0035] It can obviously be rewritten in sample form as:
[0036] This invention seeks The first derivative. Therefore, this invention has:
[0037] For W, this invention needs to solve for:
[0038]
[0039] By order: The above formula can be simplified to:
[0040] This invention can be achieved by... Solve the above equation by differentiating it and setting it to zero:
[0041] For S, the present invention needs to solve:
[0042] It can be rephrased as:
[0043] Therefore, the present invention can derive an approximate form of the above formula:
[0044] S4. Cluster the consensus representation of the optimal multi-view subspace using spectral clustering.
[0045] With the help of multi-view subspace learning, the model proposed in this invention can ultimately be expressed as:
[0046] in, This is the data matrix for the v-th view, where n is the total number of samples. Let α be the feature dimension of this data. Hyperparameters of the constraint terms. Please note, and The samples in a particular view have an overlap, making them highly coupled. Therefore, it is impossible to optimize them by using traditional alternative optimization procedures. and This invention develops an effective augmented Lagrangian method for fine-grained (ALMOND) multi-view optimization. This method cleverly bypasses the problem of uninterpretable variables inevitably introduced by previous related algorithms, and overcomes the difficulty of optimizing these two variables.
[0047] Experimental verification In this embodiment of the invention, comparative experiments were conducted with 12 state-of-the-art multi-view graph clustering methods to verify the effectiveness of the model provided by the present invention: Weighted multi-view spectral clustering (WMSC); Multi-view clustering via Cross-view Graph Diffusion (CGD); Multiple Partitions Aligned Clustering (mPAC); Multi-view Subspace Clustering via Co-training (CoMSC); Consensus One-step Multi-view Subspace Clustering (COMVSC); Co-regularized multi-view spectral clustering (Co-regularized); Fast Parameter-free Multi-view Subspace Clustering with Consensus Anchor Guidance (FPMVS); Multi-view consensus graph clustering (MCGC); Multi-view Clustering with Adaptive Neighbors (MLAN). Adaptive Neighbors); Multi-view Clustering on Topological Manifold - MVCTM; Self-weighted Multiview Clustering - SwMC This invention utilizes Yale, ORL, bbcseg13of3, Cora, Cornell, and Citeseer on benchmark datasets. Specifically: The Yale database contains raw pixel images of 15 subjects under different conditions, with each image being 64 × 64 pixels in size.
[0048] The Olivetti Research Lab's ORL face dataset consists of 400 face images across 40 different themes. For each theme, the images are described using three features: facial expression, facial details, and lighting.
[0049] bbcseg13of33 is a subset of the BBC dataset, consisting of documents from the BBC News website, corresponding to stories in five thematic areas (business, entertainment, politics, sports, and technology).
[0050] The Cora dataset includes 2,708 scientific publications, divided into 7 categories.
[0051] The Cornell dataset is a subset of the WebKB dataset collected by Cornell University, containing 195 web pages. Each web page consists of two views: content features and reference features.
[0052] Citeseer is a citation network extracted from the CiteSeer digital library, which contains 3,312 scientific publications divided into six categories.
[0053] Clustering evaluation metrics In the experiments, this invention used three clustering metrics—Normalized Mutual Information (NMI), Accuracy (ACC), Purity, and F-score—to evaluate the clustering performance. Higher values for these metrics indicate better clustering performance.
[0054] Experimental results show Figure 2-7The method of this invention outperforms its competitors in most cases. Particularly on the Cora dataset, the method of this invention shows significant advantages over the second-best method in ACC, NMI, Purity, and F-Score, respectively, by 13.79%, 16.00%, 6.76%, and 20.93%. Simultaneously, the method of this invention also achieves significant improvements on the Citeseer dataset, with improvements of 3.41%, 6.09%, 7.04%, and 5.63% in ACC, NMI, Purity, and F-Score, respectively. On the BBC dataset, improvements of 3.12%, 8.17%, 3.12%, and 5.28% are achieved relative to ACC, NMI, Purity, and F-Score, respectively. The same is true on the ORL dataset, where the method of this invention leads by 2.68%, 1.19%, 2.71%, and 4.20% across all four criteria. On the Yale dataset, MVCTM performs best in F-Score, while on the Cornell dataset, LMVSC and CoMSC slightly outperform the method of this invention in ACC and F-Score, respectively. In summary, empirical studies have demonstrated that the ALMOND method of this invention outperforms other advanced techniques in multi-view clustering.
[0055] Parameter sensitivity analysis To examine the impact of different parameter settings on clustering results, we varied the values of α, β, and λ within the ranges of [30, …, 80], [7e-6, …, 4e-5], and [700, …, 1200]. Taking the ORL dataset as an example, we can see that within a certain range, the clustering performance of different parameter settings is quite stable, such as… Figure 8 As shown. However, in our extensive experiments, we also found that clustering performance became somewhat unstable when choosing parameters over a wider range. This instability is likely due to the Lagrange parameter ρ. For simplicity, we set it to 1 from the beginning and did not search for it as a hyperparameter in our experiments. But according to the ALM principle, it grows exponentially in each iteration, i.e., ρ←1.2ρ. This may be the reason why cluster performance is somewhat unstable relative to a wider range of parameter settings.
[0056] Convergence analysis Since the optimization of this invention is essentially a non-convex problem, solved using an iterative algorithm, verifying the convergence of the proposed model is crucial. Therefore, this section empirically demonstrates the convergence and convergence speed of the algorithm. The model of this invention... Figure 9The convergence curves on the two datasets shown demonstrate the effectiveness of the optimization method of this invention. It is worth noting that although the optimization of this invention is non-convex, the model can still find the optimal solution for each variable and reach a local minimum within a small number of iterations, which verifies the effectiveness of the proposed optimization algorithm.
[0057] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0058] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0059] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0060] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
[0061] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A fine-grained multi-view image optimization method based on the augmented Lagrange method, characterized in that, The steps include the following: S1. Collect raw pixel images of multiple subjects under different conditions, and construct a multi-view data matrix based on the sample data points of the raw pixel images. The size of each raw pixel image is 64 × 64. S2. The subspace representation of the original pixel image is obtained using the subspace correlation assumption. Specifically, a consensus graph of the original pixel image is calculated based on the weights of each original pixel image in each view, thereby suppressing the inconsistencies introduced by view-wise fusion. The fine-grained fusion strategy is modeled as follows: Where i is the sample index, n is the total number of samples, and S is the consensus similarity graph. For the i-th row of S, for transpose, This is the weight matrix. For the first i The transpose of the weight vectors corresponding to each sample; For the first A tensor formed by concatenating all similar matrices on each view. for The i-th slice; Based on the augmented Lagrangian function method, the subspace representations of all original pixel images are combined into a third-order tensor, and sliced along the view direction to obtain the topological representation of each image in all views. The augmented Lagrangian function is expressed as: In the formula, For the first A data matrix for each view, where n is the total number of sample data points for each view. For the first Data dimensions of each view; Operator for calculating the sum of squares of the elements of a matrix; Let S be the similarity matrix on the v-th view; S is the consensus similarity graph. Let i be the transpose of the i-th row of the consensus similarity graph; For Lagrange multipliers; As an auxiliary variable, It is the i-th slice of a third-order tensor formed by stacking auxiliary vectors; This is the weight matrix. For the first i The transpose of the weight vectors corresponding to each sample; This is a penalty item; For hyperparameters; The topological representation of all views is as follows: In the formula, Let v be the similarity matrix on the v-th view. For the first A data matrix for each view, where n is the total number of sample data points for each view. For the first Data dimensions of each view; It is a hyperparameter; This is a penalty item; For Lagrange multipliers; As an auxiliary variable; It is the identity matrix; superscript This is the matrix transpose symbol; S3. Calculate the optimal multi-view fusion graph of the face pixel image, specifically including solving for each auxiliary variable. We obtain the optimal multi-view subspace consensus representation by calculating the weight matrix W and the consensus similarity graph S, and iteratively optimizing the objective function. S4. Cluster the consensus representation of the optimal multi-view subspace using spectral clustering.
2. The fine-grained multi-view image optimization method based on the augmented Lagrange method according to claim 1, characterized in that, In S3, each auxiliary variable is solved. The specific method is as follows: In the formula, It is the i-th slice of a third-order tensor formed by stacking auxiliary vectors; It is a hyperparameter; For the first i The weight vector corresponding to each sample for Transpose of; This is a penalty item; It is the identity matrix. Let i be the transpose of the i-th row of the consensus similarity graph; For the similarity matrix on all views The i-th slice in the spliced tensor; Let i be the i-th slice of the Lagrange multiplier.
3. The fine-grained multi-view image optimization method based on the augmented Lagrange method according to claim 2, characterized in that, The specific method for solving the weight matrix W in S3 is as follows: In the formula, It is a process variable and , for The transpose of the matrix; It is an identity matrix.
4. The fine-grained multi-view image optimization method based on the augmented Lagrange method according to claim 1, characterized in that, The solution method for the consensus similarity graph S in S3 is as follows: In the formula, The i-th row of the consensus similarity graph; λ is the hyperparameter of the sample-level fusion term, and λ is the hyperparameter of the constraint term of the consensus graph S. For the first i The transpose of the weight vector corresponding to each sample It is the i-th slice of a third-order tensor formed by stacking auxiliary vectors.
5. The fine-grained multi-view image optimization method based on the augmented Lagrange method according to claim 1, characterized in that, The clustered representation of the optimal multi-view subspace consensus representation in S4 is as follows: In the formula, Let v be the similarity matrix on the v-th view. For the similarity matrix on all views The i-th slice in the spliced tensor; S is an auxiliary variable; S is the consensus similarity graph. Let i be the transpose of the i-th row of the consensus similarity graph; This is the weight matrix. For the first i The transpose of the weight vectors corresponding to each sample; For the first A data matrix for each view, where n is the total number of sample data points for each view. For the first Data dimensions of each view; Operator for calculating the sum of squares of the elements of a matrix; Similarity matrix Hyperparameters of constraint terms; λ is the hyperparameter of the sample-level fusion term; λ is the hyperparameter of the constraint term of the consensus graph S. It is a Lagrange multiplier.