A robust multi-view image clustering method, system and device based on smooth anchor graph learning

Through the smooth anchor map learning method, Laplace matrix filtering and Frobenius norm denoising are used to generate a robust consensus anchor map, which solves the problems of high computational complexity and noise erosion in multi-view image clustering, and achieves efficient and accurate image clustering.

CN119963867BActive Publication Date: 2025-08-12XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510064101.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-08-12
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

When facing large-scale data and complex noise erosion, the existing multi-view image clustering method has high computational complexity and low clustering accuracy, making it difficult to effectively process label-free multi-view image data.

Method used

Using a method based on smooth anchor map learning, the Laplace matrix filtering denoising, generate anchor point sets and adaptively learn anchor point maps, combined with local popular learning and Frobenius norm denoising, a robust consensus anchor point map is obtained to improve clustering efficiency and accuracy.

Benefits of technology

It significantly improves clustering efficiency and accuracy in complex noise environments, can effectively process large-scale label-free multi-view image data, and improves the robustness of multi-view clustering tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963867B_ABST
    Figure CN119963867B_ABST
Patent Text Reader

Abstract

The present invention is aimed at large-scale unlabeled image data under complex noise erosion, and discloses a robust multi-view image clustering method, system and device based on smoothed anchor graph learning, which belongs to the field of information technology. The method includes: 1. Acquiring large-scale multi-view image features eroded by noise; 2. Performing graph filtering smoothing and denoising on the acquired image features; 3. Generating a representative anchor set, and adaptively learning the anchor graph between each view anchor set and the smoothed image features; 4. Applying a local popular learning method to fuse the anchor graphs of different views to obtain a weighted anchor graph; 5. Smoothing and denoising the weighted anchor graph under the Frobenius norm to obtain a consensus anchor graph; 6. Obtaining a consensus spectrum embedding and the corresponding clustering results. When processing the clustering task of real large-scale multi-view image data containing complex noise, the method of the present invention can significantly improve the clustering efficiency and the robustness to complex noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image clustering analysis in machine learning, and in particular to a robust multi-view image clustering method, system and device based on smooth anchor graph learning. Background Art

[0002] In recent years, with the rapid development of information technology and the Internet, multi-view image data has experienced explosive growth. In order to process these unlabeled multi-view image data, multi-view image clustering analysis technology, which can directly cluster unlabeled multi-view image data, has been rapidly developed in recent years.

[0003] However, existing multi-view image clustering methods still face the following two challenges: 1) With the explosive growth of data volumes, traditional multi-view image clustering methods are computationally complex and therefore struggle to meet the requirements for efficient image clustering. 2) Real-world data is often subject to complex noise due to factors such as sensor errors, data transmission interference, changing environmental conditions, or camera failures. Traditional multi-view image clustering methods struggle to maintain high clustering accuracy in the face of complex noise.

[0004] In summary, the related technologies are difficult to process multi-view image data with increasing size and eroded by complex noise. Therefore, the problems existing in the related technologies need to be solved urgently. Summary of the Invention

[0005] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a robust multi-view image clustering method based on smooth anchor graph learning. The method reduces the computational complexity when facing large-scale image data by learning weighted anchor graphs of all views, and performs graph filtering on the noise-eroded image features using the Laplacian matrix to alleviate the complex noise erosion in the original image features. At the same time, a consensus anchor graph is obtained by smoothing and denoising the weighted anchor graph under the Frobenius norm metric to cope with the erosion of the anchor graph by complex noise, which greatly improves the efficiency and robustness when processing large-scale image data eroded by complex noise.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] In one aspect, the present invention provides a robust multi-view image clustering method based on smooth anchor graph learning, the method comprising:

[0008] Acquire large-scale multi-view image features eroded by complex noise;

[0009] Performing graph filtering on the obtained image features to smooth and denoise;

[0010] Generate a representative anchor point set and adaptively learn the anchor point map between each view anchor point set and the smoothed image data;

[0011] Apply local manifold learning method to fuse anchor maps of different views to obtain weighted anchor maps;

[0012] The weighted anchor graph is smoothed and denoised under the Frobenius norm to obtain a consensus anchor graph;

[0013] Obtain consensus spectrum embedding and corresponding clustering results.

[0014] Optionally, obtaining large-scale multi-view image features corroded by complex noise includes:

[0015] Multiple salient features are extracted from visible light data, infrared data and SAR imaging data corroded by complex noise to obtain large-scale multi-view image features.

[0016] Optionally, performing graph filtering on the obtained image features to perform smoothing and denoising includes:

[0017] A Laplacian matrix is constructed independently for each view, and the Laplacian matrix is used to perform graph filtering on the original image features eroded by complex noise to alleviate the complex noise erosion of the original image features:

[0018]

[0019] where X (v) represents the original image features of the v-th view, L (v) represents the Laplacian matrix of the v-th view, Represents the image features after smoothing by the Laplace matrix, I represents the unit matrix, and μ represents a balance parameter.

[0020] Optionally, a representative anchor point set is generated, and an anchor point map between each view anchor point set and the smoothed image data is adaptively learned, including:

[0021] For the obtained smoothed image features containing V views eroded by noise Among them, n is the number of images that need to be clustered, d (v) is the dimension of the image features contained in the vth view; the anchor set is generated independently for all views Among them, m is the number of anchor points; then the anchor map of each view is adaptively learned based on the topological relationship between the anchor point set and the smoothed image features.

[0022] Optionally, the step of adaptively learning the anchor point graph of each view according to the topological relationship between the anchor point set and the smoothed image features includes:

[0023]

[0024] in represents the element in row i and column j of the anchor graph of view v, Φ(·) represents the distance metric function, and k represents that the anchor graph only records the features of the i-th original image. The relationship between the k nearest anchor points; Represents the original image feature of the i-th The jth anchor point closest to it.

[0025] Optionally, a local manifold learning method is applied to fuse the anchor maps of different views to obtain a weighted anchor map, including:

[0026] Apply the local popular learning method to assign different weights to the anchor maps of different views, and fuse the anchor maps of different views to obtain the weighted anchor map:

[0027]

[0028] where α (v) and B (v) Represent the weight of the v-th anchor graph and the v-th anchor graph respectively. By performing weighted summation on the anchor graphs from different views, the weighted anchor graph B is obtained. A .

[0029] Optionally, the weighted anchor graph is smoothed and denoised under the Frobenius norm to obtain a consensus anchor graph, including:

[0030] Under the Frobenius norm metric, the influence of noise and outliers in the weighted anchor graph is eliminated to obtain a robust consensus anchor graph:

[0031]

[0032] in represents the learned consensus anchor graph, Represents the weighted anchor map obtained by weighted summation of anchor maps from different views.

[0033] Optionally, obtain the consensus spectrum embedding and corresponding clustering results, including:

[0034] Learn the consensus spectral embedding based on the consensus anchor graph, and learn the final cluster indicator matrix based on the consensus spectral embedding:

[0035]

[0036] where Tr[·] represents the trace of the matrix, F is the consensus spectral embedding representation of the learned consensus anchor graph, and I represents the identity matrix. and represent the transpose of F and Z respectively; finally, the clustering result is obtained by performing K-means clustering on the consensus spectral embedding representation.

[0037] Another aspect of the present invention further provides a system for implementing the robust multi-view image clustering method based on smooth anchor graph learning, comprising:

[0038] Feature extraction module: obtains multi-view image features with complex noise erosion for clustering tasks;

[0039] Laplacian filter module: constructs the Laplacian matrix to obtain image features, and uses the Laplacian matrix to perform graph filtering on the original image features to reduce the complex noise erosion in the original image features;

[0040] Anchor map generation module: First, it generates an anchor point set for each view. Then, based on the topological relationship between the anchor point set and the smoothed image features, it adaptively generates an anchor map corresponding to each view.

[0041] Anchor graph weighting module: This module uses local popular learning to assign different weights to the anchor graphs of all views and fuses the anchor graphs of different views to obtain a weighted anchor graph.

[0042] Anchor graph smoothing module: This module smoothes the weighted anchor graph using the Frobenius norm metric to reduce the impact of noise in the weighted anchor graph and obtain a robust consensus anchor graph.

[0043] Clustering result acquisition module: obtains consensus spectrum embedding according to the consensus anchor graph, and obtains the final clustering result through the K-means clustering method.

[0044] In another aspect of the present invention, a multi-view image clustering device is provided, which includes a memory and a processor, wherein the memory is used to store instructions and data, and the processor is used to execute instructions stored in the memory; the instructions are used to implement the robust multi-view image clustering method based on smooth anchor graph learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical implementation schemes in the embodiments of this application, this application provides relevant drawings. It should be noted that the drawings described in this section are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on the present invention and the drawings without inventive work.

[0046] Figure 1 is a flow chart of the algorithm of the present invention;

[0047] Figure 2 is the optimization iteration of the objective function;

[0048] Figure 3 This is the comparison of clustering robustness. DETAILED DESCRIPTION

[0049] In order to describe the implementation process of the present invention in detail, the following will be described in more detail with reference to the relevant drawings. It should be noted that the implementation examples described herein are only used to better illustrate the present invention, and the present invention can be implemented in various forms and is not limited to the examples described herein.

[0050] The specific implementation steps of the present invention are as follows Figure 1 As shown, the steps are as follows:

[0051] Step S01: Acquire large-scale multi-view image features eroded by complex noise. It should be noted that in some scenarios, the available data only contains imageable data, and different saliency feature extraction methods are required for the imageable data to obtain multi-view image features. For example, when the object only contains visible light data, multiple saliency feature extraction methods can be used to extract features to obtain multi-view image features. Noise includes salt and pepper noise, Gaussian noise, Poisson noise, speckle noise, and other noise caused by sensor errors, data transmission interference, changes in environmental conditions, or camera equipment failure, among other factors.

[0052] Step S02: For the obtained multi-view image features containing V views Generate Laplacian matrices for all views independently Use the Laplacian matrix to perform graph filtering on the noise-eroded original image features to alleviate the complex noise erosion in the original image features:

[0053]

[0054] where X (v) represents the original image features of the v-th view, L (v) represents the Laplacian matrix of the v-th view, Represents the image features after smoothing by the Laplace matrix, I represents the unit matrix, and μ represents a balance parameter.

[0055] Step S03: Generate a representative anchor point set, and then use adaptive learning to generate an anchor point map between the anchor point set and the smoothed filtered image features for each view. Among them, n is the number of images that need to be clustered, d (v) is the dimension of the image features contained in the vth view; the anchor set is generated independently for all views Where m is the number of anchor points. Then, the anchor graph of each view is adaptively learned based on the topological relationship between the anchor point set and the image features after smoothing filtering.

[0056] For the process of generating anchor point sets for all views, methods including random sampling, K-means clustering, hierarchical clustering, etc. can be used. In some embodiments, using K-means clustering to generate anchor point sets as an example, for all views v∈V, the anchor point set of view v can be generated by the following method:

[0057]

[0058] Where m represents the number of anchor points, represents the jth anchor point, represents the i-th original image feature, represents the region formed by the j-th anchor point as a cluster, It should be noted that for the present invention, the number of anchor points between different views should be unified to ensure that the dimensions of the anchor maps of different views are consistent.

[0059] For the process of generating anchor graphs for all views, various anchor graph forms can be generated, such as fully connected anchor graphs, k-nearest neighbor anchor graphs, etc. In some embodiments, a k-nearest neighbor anchor graph is generated. For all views v∈V, the anchor graph of view v can be generated by the following method:

[0060]

[0061] in represents the element in row i and column j of the anchor graph of view v, Φ(·) represents the distance metric function, and k represents that the anchor graph only records the features of the i-th original image. The relationship between the k nearest anchor points, Represents the original image feature of the i-th The jth anchor point closest to it.

[0062] Step S04: Apply the local manifold learning method to fuse the anchor maps of different views. In some embodiments, different weights are assigned to the anchor maps of different views, and the fused anchor map is regarded as the weighted sum of the anchor maps of different views to obtain a weighted anchor map:

[0063]

[0064] where α (v) and B (v) Represent the weight of the v-th anchor graph and the v-th anchor graph respectively. By performing weighted summation on the anchor graphs from different views, the weighted anchor graph B can be obtained. A .

[0065] Step S05: Smoothing and denoising the weighted anchor graph under the Frobenius norm metric to obtain a consensus anchor graph, so as to filter out the complex noise contained in the real world.

[0066] In some embodiments, the similarity in the anchor feature space is measured under the Frobenius norm:

[0067]

[0068] in represents the learned consensus anchor graph, Represents the weighted anchor map obtained by weighted summation of anchor maps from different views.

[0069] Step S06: Learn the consensus spectral embedding through the consensus anchor graph, and finally perform K-means clustering on the consensus spectral embedding to obtain the final clustering result.

[0070] In some embodiments, a method for learning spectral embedding is performed using the normalized cut (Ncut) theory, and the specific process is as follows:

[0071]

[0072] where Tr[·] represents the trace of the matrix, I represents the identity matrix, and F is the consensus spectral embedding of the learned consensus anchor graph. and Denote the transpose of F and Z respectively. By performing This is equivalent to obtaining a full sample graph between all samples by executing This is equivalent to obtaining the relaxed Laplace matrix of the entire sample graph.

[0073] The elements in the consensus spectrum embedding F are distributed in the entire real number space, and it is difficult to directly obtain the clustering result. In some embodiments, the K-means clustering method can be used to obtain the categories corresponding to the samples for the obtained consensus spectrum embedding F to obtain the final clustering result. The iterative convergence of the method of the present invention is as follows Figure 2 As shown in the figure, it can be seen that this method can converge within 15 iterations and has a good convergence speed.

[0074] The embodiments of the present invention can effectively process large-scale, unlabeled multi-view image data under complex noise erosion, significantly improving the efficiency and accuracy of multi-view clustering tasks. Compared with existing multi-view clustering methods, the clustering accuracy of the present invention is higher, as shown in Table 1.

[0075] Table 1 Comparison of clustering results between the proposed method and advanced multi-view clustering methods on real datasets

[0076]

[0077]

[0078] Table 1 shows the clustering accuracy comparison of the present invention and other advanced multi-view clustering methods on four image datasets. NUSW is a large-scale dataset with more than 10,000 samples. The results show that the clustering effect of the present invention has the highest or second highest clustering accuracy among different clustering indicators. Compared with the multi-view clustering method in the prior art, the algorithm of the present invention is more robust when facing complex noise erosion, such as Figure 3 shown. Figure 3 The robustness of clustering of the present invention and other advanced clustering methods on the image dataset AWA is compared. The results show that under different degrees of salt and pepper noise erosion, the present invention has stable clustering accuracy and exhibits excellent robustness.

[0079] On the other hand, the present invention also provides a robust multi-view image clustering system based on smooth anchor graph learning, comprising:

[0080] Feature extraction module: obtains multi-view image features with complex noise erosion for clustering tasks;

[0081] Laplacian filter module: constructs the Laplacian matrix to obtain image features, and uses the Laplacian matrix to perform graph filtering on the original image features to reduce the complex noise erosion in the original image features;

[0082] Anchor map generation module: First, it generates an anchor point set for each view. Then, based on the topological relationship between the anchor point set and the smoothed image features, it adaptively generates an anchor map corresponding to each view.

[0083] Anchor graph weighting module: This module uses local popular learning to assign different weights to the anchor graphs of all views and fuses the anchor graphs of different views to obtain a weighted anchor graph.

[0084] Anchor graph smoothing module: This module smoothes the weighted anchor graph using the Frobenius norm metric to reduce the impact of noise in the weighted anchor graph and obtain a robust consensus anchor graph.

[0085] Clustering result acquisition module: obtains consensus spectrum embedding according to the consensus anchor graph, and obtains the final clustering result through the K-means clustering method.

[0086] In another aspect, the present invention further provides a multi-view image clustering device comprising a memory and a processor. The memory is configured to store instructions and data, and the processor is configured to execute instructions stored in the memory; the instructions are configured to implement the robust multi-view image clustering method based on smooth anchor graph learning.

[0087] It should be noted that although a detailed implementation example has been provided to help those skilled in the art understand the overall framework and technical details of the present invention, the present invention is not limited to the described embodiment. Those skilled in the art may make equivalent modifications or substitutions while adhering to the principles and spirit of the present invention, and such equivalent modifications or substitutions are all within the scope of the claims of the present invention.

Claims

1. A robust multi-view image clustering method based on smooth anchor graph learning, characterized by: include: Acquire large-scale multi-view image features eroded by complex noise; Performing graph filtering on the obtained image features to smooth and denoise; Generate a representative anchor point set and adaptively learn the anchor point map between each view anchor point set and the smoothed image data; Apply local manifold learning method to fuse anchor maps of different views to obtain weighted anchor maps; The weighted anchor graph is smoothed and denoised under the Frobenius norm to obtain a consensus anchor graph; Obtain consensus spectrum embedding and corresponding clustering results; The weighted anchor graph is smoothed and denoised under the Frobenius norm to obtain a consensus anchor graph, including: Under the Frobenius norm metric, the influence of noise and outliers in the weighted anchor graph is eliminated to obtain a robust consensus anchor graph: in represents the learned consensus anchor graph, represents the weighted anchor map obtained by weighted summation of anchor maps of different views; Obtain consensus spectrum embedding and corresponding clustering results, including: Learn the consensus spectral embedding based on the consensus anchor graph, and learn the final cluster indicator matrix based on the consensus spectral embedding: where Tr[·] represents the trace of the matrix, F is the consensus spectral embedding representation of the learned consensus anchor graph, and I represents the identity matrix. and represent the transpose of F and Z respectively; finally, the clustering result is obtained by performing K-means clustering on the consensus spectral embedding representation.

2. The robust multi-view image clustering method based on smooth anchor graph learning according to claim 1, characterized in that The method of obtaining large-scale multi-view image features corroded by complex noise includes: Multiple salient features are extracted from visible light data, infrared data and SAR imaging data corroded by complex noise to obtain large-scale multi-view image features.

3. The robust multi-view image clustering method based on smooth anchor graph learning according to claim 1, characterized in that The performing graph filtering on the obtained image features to perform smoothing and denoising includes: A Laplacian matrix is constructed independently for each view, and the Laplacian matrix is used to perform graph filtering on the original image features eroded by complex noise to alleviate the complex noise erosion of the original image features: where X (v) represents the original image features of the v-th view, L (v) represents the Laplacian matrix of the v-th view, Represents the image features after smoothing by the Laplace matrix, I represents the unit matrix, and μ represents a balance parameter.

4. The robust multi-view image clustering method based on smooth anchor graph learning according to claim 1, characterized in that Generate a representative anchor point set and adaptively learn the anchor point map between each view anchor point set and the smoothed image data, including: For the obtained smoothed image features containing V views eroded by noise Among them, n is the number of images that need to be clustered, d (v) is the dimension of the image features contained in the vth view; the anchor set is generated independently for all views Among them, m is the number of anchor points; then the anchor map of each view is adaptively learned based on the topological relationship between the anchor point set and the smoothed image features.

5. The robust multi-view image clustering method based on smooth anchor graph learning according to claim 4, characterized in that The steps of adaptively learning the anchor graph of each view according to the topological relationship between the anchor set and the smoothed image features include: in represents the element in row i and column j of the anchor graph of view v, Φ(·) represents the distance metric function, and k represents that the anchor graph only records the features of the i-th original image. The relationship between the k nearest anchor points; Represents the original image feature of the i-th The jth anchor point closest to it.

6. The robust multi-view image clustering method based on smooth anchor graph learning according to claim 1, characterized in that Apply the local manifold learning method to fuse the anchor maps of different views to obtain a weighted anchor map, including: Apply the local popular learning method to assign different weights to the anchor maps of different views, and fuse the anchor maps of different views to obtain the weighted anchor map: where α (v) and B (v) Represent the weight of the v-th anchor graph and the v-th anchor graph respectively. By performing weighted summation on the anchor graphs from different views, the weighted anchor graph B is obtained. A .

7. A system for implementing the robust multi-view image clustering method based on smooth anchor graph learning according to any one of claims 1 to 6, characterized in that: include: Feature extraction module: obtains multi-view image features with complex noise erosion for clustering tasks; Laplacian filter module: constructs the Laplacian matrix to obtain image features, and uses the Laplacian matrix to perform graph filtering on the original image features to reduce the complex noise erosion in the original image features; Anchor map generation module: First, an anchor set is generated for each view. Then, based on the topological relationship between the anchor set and the smoothed large-scale multi-view image features, an anchor map corresponding to each view is adaptively generated. Anchor graph weighting module: This module uses local popular learning to assign different weights to the anchor graphs of all views and fuses the anchor graphs of different views to obtain a weighted anchor graph. Anchor graph smoothing module: This module smoothes the weighted anchor graph using the Frobenius norm metric to reduce the impact of noise in the weighted anchor graph and obtain a robust consensus anchor graph. Clustering result acquisition module: obtains consensus spectrum embedding based on the consensus anchor graph, and obtains the final clustering result through the K-means clustering method.

8. A multi-view image clustering device, characterized in that: The multi-view image clustering device includes a memory and a processor, wherein the memory is used to store instructions and data, and the processor is used to execute instructions stored in the memory; wherein the instructions are used to implement the robust multi-view image clustering method based on smooth anchor graph learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Anchor point-based mineral identification and clustering analysis method

    CN118298411A

  • Multi-view clustering method and system based on matrix decomposition and multi-partition alignment

    US20240111829A1