An incomplete multi-view clustering system and method based on confidence map and tensor feature selection

CN122761003APending Publication Date: 2026-09-15GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610944586.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-15

Smart Images

  • Figure CN122761003A_ABST
    Figure CN122761003A_ABST
Patent Text Reader

Abstract

The application discloses an incomplete multi-view clustering system and method based on a confidence graph and tensor feature selection, and the system comprises the following modules: a data preprocessing and local graph initialization module, which is used for constructing an initial similarity graph, a confidence graph and a complete similarity graph to be learned; a confidence-guided graph dynamic reconstruction module, which is used for iteratively updating and reconstructing the complete similarity graph; a tensor decoupling and high-order relationship extraction module, which is used for extracting global high-order consistency manifold structures across views; a feature selection and clustering module, which is used for outputting clustering results; and a closed-loop optimization module, which is used for integrating optimization objectives of the confidence-guided graph dynamic reconstruction module, the tensor decoupling and high-order relationship extraction module and the feature selection and clustering module into a total objective function and iteratively solving the total objective function. The application realizes perfect balance between local manifold and global structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data mining and unsupervised learning technology, and in particular to an incomplete multi-view clustering system and method based on confidence graphs and tensor feature selection. Background Technology

[0002] Clustering is a fundamental unsupervised learning technique for discovering latent patterns in complex data. In the era of big data, objects are often represented by multiple modalities or views (such as color, texture, shape, etc.). Multi-view clustering (MVC) improves clustering performance by leveraging complementary information between views. However, in practical applications, multi-view data is often incomplete due to sensor failures or storage limitations, leading to the incomplete multi-view clustering (IMC) problem.

[0003] Existing incomplete multi-view clustering methods and traditional unsupervised feature selection (UFS) techniques face the following three main drawbacks when processing this type of data: 1. Rigid similarity matrix updates and poor noise resistance: Existing UFS modules are highly dependent on the quality of the initial graph. During iterative updates, the initial similarity matrix remains static, which cannot effectively suppress high-dimensional noise and interference from outliers (discrete points) introduced by incomplete data and sensors, leading to the propagation of pseudo-topological connections during data recovery.

[0004] 2. Ignoring correlations between higher-order views: Most existing methods only deal with pairwise or low-order relationships between views, failing to fully explore and capture global higher-order tensor relationships across views.

[0005] 3. Imbalance between local and global relationships: Existing feature extraction frameworks lack the synergistic effect between the manifold structure within the local view and the globally consistent manifold structure, and cannot take into account both local denoising and global structure alignment. Summary of the Invention

[0006] This invention aims to solve the technical problems existing in incomplete multi-view clustering methods, such as the similarity matrix not being dynamically updated, susceptibility to noise and discrete point interference, neglect of higher-order relationships across views, and imbalance between local and global structures. It provides an incomplete multi-view clustering system and method based on confidence graphs and tensor feature selection.

[0007] On the one hand, to achieve the above objectives, the present invention provides an incomplete multi-view clustering system based on confidence graphs and tensor feature selection, comprising: The data preprocessing and local graph initialization module is used to extract the available data matrix under each view for incomplete multi-view data containing missing instances, and to construct an initial similarity map, confidence map and complete similarity map to be learned based on the available data matrix. The confidence-guided graph dynamic reconstruction module is used to iteratively update and reconstruct the complete similarity graph based on a self-representation learning mechanism, combined with the confidence graph as a structural filter. The tensor decoupling and higher-order relation extraction module is used to stack the complete similarity graphs after the reconstruction of each view to construct a self-representation tensor and extract the global higher-order consistent manifold structure across views. The feature selection and clustering module is used to generate a globally consistent Laplacian matrix based on the global high-order consistent manifold structure, and to establish a feature selection optimization objective by combining the local manifold distribution of the data. It also filters discriminative feature subsets through orthogonal basis decomposition and sparse constraints, and outputs clustering results. The closed-loop optimization module is used to integrate the optimization objectives of the confidence-guided graph dynamic reconstruction module, tensor decoupling and higher-order relation extraction module, and feature selection and clustering module into a general objective function, and solve iteratively.

[0008] Preferably, constructing the confidence graph includes: Based on the available data matrix, a binary indicator matrix is ​​constructed using K-nearest neighbor relationships, and the inner product of the binary indicator matrix is ​​calculated to obtain a matrix of common neighbors. The common neighbor number matrix is ​​percentage-normalized to generate the confidence graph, which is used to quantify the reliability of local connections between samples.

[0009] Preferably, the confidence-guided graph dynamic reconstruction module utilizes the confidence graph as a structure filter in the following ways: In the objective function of reconstructing the complete similarity map, the Hadamard product operator is introduced to perform element-wise multiplication between the complete similarity map to be learned and the aligned confidence map to form a mask operation, which forcibly reduces the weights of the connections that are judged as low reliability by the confidence map.

[0010] Preferably, extracting the global higher-order consistent manifold structure across views includes: The self-representation tensor is decoupled and decomposed into an essential tensor and a noise tensor. A weighted tensor kernel norm constraint is applied to the essential tensor, and the global high-order consistent manifold structure across the view is extracted.

[0011] Preferably, a weighted tensor nuclear norm constraint is applied to the essential tensor, specifically as follows: ; In the formula, It is an essential tensor; Represents the number of views; This represents the total number of samples actually processed. For the essential tensor in the Fourier transform domain, the first... One slice; The essential tensor in the Fourier transform domain is the first... The first slice Large singular values; These are the weighting coefficients; Let represent the weighted tensor nuclear norm.

[0012] Preferably, the feature selection optimization objective established by the feature selection and clustering module is: ; In the formula, For the complete data matrix; This refers to the feature projection matrix or the feature weight matrix; It is an orthogonal basis matrix; This is a pseudo-clustering label matrix; for Norm constraints; It is a globally consistent Laplace matrix; It is the Frobenius norm; , All are balance parameters; Let be the trace of the matrix.

[0013] On the other hand, to achieve the above objectives, the present invention also provides an incomplete multi-view clustering method based on confidence graphs and tensor feature selection, comprising: S1. Obtain incomplete multi-view data and construct an initial similarity map and confidence map based on the available data in each view; S2. Construct the first optimization subproblem containing self-representation reconstruction terms and confidence graph filtering terms, introduce a data completion matrix to fill in missing instances, and iteratively update the complete similarity graph of each view; S3. Stack the updated complete similarity maps of each view into a self-representation tensor, construct a second optimization subproblem containing tensor low-rank constraints, decompose the self-representation tensor into an essential tensor and a noise tensor, and apply a weighted tensor nuclear norm constraint to the essential tensor to capture the high-order correlation between views. S4. Construct a global Laplacian matrix based on the essential tensor, and construct a third optimization subproblem that includes feature projection, orthogonal basis clustering and local manifold preservation, and solve for the feature projection matrix; S5. Integrate the first optimization sub-problem, the second optimization sub-problem, and the third optimization sub-problem into a general objective function, perform joint optimization using the alternating direction multiplier method, and feed back the solution result of the feature projection matrix to S2 during the iteration process to dynamically adjust the update direction of the similarity map.

[0014] Preferably, it further includes: After the iteration converges, the original features are sorted according to the row vector norm of the feature projection matrix, a subset of discriminative features is selected, and the clustering algorithm is input to output the final clustering result.

[0015] Compared with the prior art, the present invention has the following advantages and technical effects: (1) Improved the noise resistance and robustness of the model: This invention integrates confidence graphs and unsupervised feature selection for the first time. The confidence graphs, as structural filters, effectively prevent the propagation of pseudo-topological connections during data recovery and significantly reduce the interference of discrete points and high-order noise.

[0016] (2) Fully captures and utilizes high-order relationships across views: This invention processes the stacked complete similarity map using weighted tensor kernel norm (WTNN), breaking the limitation of traditional methods that only consider low-order correlations, and can adaptively distill global consistency information across views.

[0017] (3) Achieving a perfect balance between local manifold and global structure: This invention combines the global high-order relations and feature selection obtained by WTNN with the local manifold structure preserved by the clustering module. The purified features can iteratively optimize the construction of the graph, forming a dynamic closed-loop feedback system that ensures the discriminative power of the features while taking into account local denoising. Even with a data missing rate of up to 40% and heterogeneous feature distribution, it can still maintain extremely high clustering accuracy (ACC), normalized mutual information (NMI) and purity (PUR). Attached Figure Description

[0018] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of an incomplete multi-view clustering method based on confidence graph and tensor feature selection according to an embodiment of the present invention. Detailed Implementation

[0019] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0021] This embodiment proposes an incomplete multi-view clustering system based on confidence graphs and tensor feature selection, including: The data preprocessing and local graph initialization module is used to extract the available data matrix under each view for incomplete multi-view data containing missing instances, and to construct an initial similarity map, confidence map and complete similarity map to be learned based on the available data matrix. The confidence-guided graph dynamic reconstruction module is used to iteratively update and reconstruct the complete similarity graph based on a self-representation learning mechanism, combined with the confidence graph as a structural filter. The tensor decoupling and higher-order relation extraction module is used to stack the complete similarity graphs after the reconstruction of each view to construct a self-representation tensor and extract the global higher-order consistent manifold structure across views. The feature selection and clustering module is used to generate a globally consistent Laplacian matrix based on the global high-order consistent manifold structure, and to establish a feature selection optimization objective by combining the local manifold distribution of the data. It also filters discriminative feature subsets through orthogonal basis decomposition and sparse constraints, and outputs clustering results. The closed-loop optimization module is used to integrate the optimization objectives of the confidence-guided graph dynamic reconstruction module, tensor decoupling and higher-order relation extraction module, and feature selection and clustering module into a general objective function, and solve iteratively.

[0022] Furthermore, the processing steps of the data preprocessing and local graph initialization module include: For incomplete multi-view data containing missing instances, the available data matrix for each view is extracted. Based on the extracted available data, an initial similarity map between available samples is calculated using Euclidean distance; simultaneously, based on the available data... A binary indicator matrix is ​​constructed based on nearest neighbor relationships. A confidence graph is generated by calculating the number of common neighbors and performing percentage standardization. This confidence graph is used to quantify the reliability of local connections between samples.

[0023] Specifically, input an incomplete multi-view data matrix. Extract the available multi-view data matrix .

[0024] Initialize alignment matrix (Indicating the location of missing samples) and (Indicates the location of available samples).

[0025] Furthermore, constructing the confidence graph includes: Based on the available data matrix, a binary indicator matrix is ​​constructed using K-nearest neighbor relationships, and the inner product of the binary indicator matrix is ​​calculated to obtain a matrix of common neighbors. The common neighbor number matrix is ​​percentage-normalized to generate the confidence graph, which is used to quantify the reliability of local connections between samples.

[0026] Specifically, extract the available multi-view data matrix from each view. First, the similarity graph between available samples is calculated using Euclidean distance. : ; In the formula, Representing similarity graphs The row and column indices of the sample (i.e., the first row in the dataset) The and the first (one sample data point) These are usable physical samples, representing fragments of multimodal data that were actually successfully acquired. The square of the Euclidean distance is used to physically measure the absolute difference between two samples in the original sensor feature space. The smaller the distance, the more similar the two samples are in the physical world (e.g., two face images are extremely similar, or two texts have highly overlapping features). Given the number of available samples, and to eliminate sensor noise and interference from discrete points, a binary nearest neighbor matrix is ​​constructed based on the available data. The specific calculations are as follows: ; In the formula, Indicates sample The set of k-nearest neighbors; Indicates sample The set of k-nearest neighbors; these two conditions are used together to rigorously determine whether samples have an exact adjacency relationship in the local topology. Represents the nearest neighbor region, physical space, or feature space. K Nearest neighbor set; It is a binary nearest neighbor matrix, which, in a practical physical sense, is used to force a determination of whether two samples have a precise adjacency relationship in terms of local features.

[0027] Then calculate the confidence matrix. ; The final confidence plot is obtained by standardizing the maximum value by percentage. : ; In the formula, These are usable physical samples, representing fragments of multimodal data that were actually successfully acquired. The confidence matrix or common neighbor count matrix is ​​derived from nearest neighbor clustering, and its elements are calculated by the inner product of the binary indicator matrix vectors. The quotient of the two values ​​represents the maximum value in the matrix, thus achieving percentage normalization of the data and minimizing the interference of noise and discrete points on the construction of similarity relationships. The confidence graph matrix is ​​not a simple distance metric, but rather a measure of the number of common neighbors. The inner product form of the algorithm is used to quantify the "reliability percentage" of connections between samples. In practical applications, it can effectively identify pseudo-connections that appear similar due to accidental noise, providing high-confidence structural prior information for subsequent data recovery.

[0028] In this embodiment, when constructing the similarity map of available samples, in addition to using Euclidean distance, Gaussian kernel distance can also be used as an alternative calculation under a specific smooth data distribution.

[0029] Furthermore, the processing procedure of the confidence-guided graph dynamic reconstruction module includes: To extend the locally available view to the global view, this embodiment introduces a self-representation mechanism, which combines a missing data imputation matrix to establish a self-representation reconstruction model.

[0030] Its core self-representation equation is: ; Iteratively update the complete similarity map At that time, using confidence graphs As a structure filter, its corresponding objective function is: ; In the formula, To complete the matrix, For the alignment matrix, Used to pinpoint the exact location of missing physical location data (0-1 matrix). These are the feature values ​​to be filled, dynamically estimated by the algorithm. Together, they constitute the complete set of physical signals after the repair. For self-representation coefficient matrix / complete similarity graph, based on self-representation theory, any sample in nature or system can be approximated by a linear combination of other samples of the same kind; It is about learning the global similarity weights. This is a self-representation error matrix, representing objectively existing and unavoidable system measurement errors or random noise. This represents the Hadamard product / elemental multiplication, and the physical meaning of this operator lies in the "mask operation"; it involves comparing the reconstructed similarity matrix with the confidence graph. (through After alignment, element-wise multiplication forces the filtering out of discrete points with extremely low confidence and noise connections, ensuring that the recovered data strictly follows a reliable topology. As a balancing parameter, the weights of the confidence map mask operation (structural filter) of the samples can be used to adjust the reconstruction of the similarity matrix, so as to effectively eliminate the influence of pseudo-topological connections caused by discrete points and noise.

[0031] Furthermore, the processing steps of the tensor decoupling and higher-order relation extraction module include: The reconstructed similarity maps of each view are stacked and rotated along the view dimension to construct a self-representation tensor. .

[0032] To effectively separate valid information from noise, the self-representation tensor is decomposed into the essential tensor. and noise tensor By imposing weighted tensor nuclear norm constraints on the essential tensor, the global higher-order consistent manifold structure across the view is extracted.

[0033] Specifically, the complete similarity map after reconstructing each view. By stacking and rotating along the view dimension, a representation tensor is constructed. To effectively separate valid information from noise, the self-representation tensor is decomposed into the essential tensor. and noise tensor .

[0034] We use the weighted tensor kernel norm (WTNN) for low-rank constraint optimization of the essential tensor. The core calculation formula is as follows: ; In the formula, The essential tensor represents the most essential network of all samples in the objective physical world under various views after removing residual noise. This represents the number of views. In practical applications, it can represent different sensor types that capture the same object, such as image features and text features. This represents the total number of samples actually processed. For the essential tensor in the Fourier transform domain, the first... One slice; The essential tensor in the Fourier transform domain is the first... The first slice Large singular values, in their physical sense, represent the magnitude of principal component energy after data from different modes are aligned. These are the weighting coefficients; The weighted tensor nuclear norm (WTNN) is used to assign a set of weight coefficients to different singular values ​​of tensor slices, making the calculation of the nuclear norm more flexible and thus effectively extracting the high-order globally consistent manifold structure of multi-view data.

[0035] Furthermore, the processing steps of the feature selection and clustering module include: After obtaining the global consistency relation optimized by tensor low rank, this is used to guide unsupervised feature selection. This module decomposes the target matrix into potential cluster center indicators and sparse representations of different classes through orthogonal basis decomposition. Its core objective function (sub-terms) for optimization is as follows: ; In the formula, For the complete data matrix, that is This represents the complete physical sample feature data after filling in the missing values. The feature projection matrix or feature weight matrix has the actual physical meaning of measuring the contribution of each feature in the original data (e.g., removing useless background pixels in an image, or discarding invalid sensor dimensions). It is an orthogonal basis matrix; This is a pseudo-clustering label matrix. Used to project high-dimensional physical data into a low-dimensional cluster space; This indicates the probability that a sample ultimately belongs to each cluster category, through... A low-dimensional representation that approximates the original data. for Norm constraints are key to feature selection. In actual computation, this constraint causes row sparsity in the feature projection matrix, thereby automatically masking redundant and irrelevant sensor features and retaining the most discriminative core data. It is a globally consistent Laplace matrix. ,in, It is the final high-confidence similarity map that integrates information from multiple views. The trace at the end of the formula... The regularization term ensures that two samples that are originally highly similar in the objective world (i.e., in...) are similar to each other. Even after feature selection and dimensionality reduction, the feature representations of sample pairs with high weights remain similar, achieving a perfect combination of local manifold structure and global topology. The Frobenius norm is primarily used to measure the error or fit between matrices. These are balancing parameters used for control. norm ( The penalty weights of the feature projection matrix induce row sparsity, thereby extracting the main features of orthogonal basis clustering of local structures and filtering redundant features. The balancing parameter is used to control the weight of the local structure preservation term, ensuring that the local structure information of the original complete data is still preserved after feature dimensionality reduction; The trace of the matrix is ​​used to calculate the sum of the diagonal elements of the square matrix. In the model, it is used to construct and solve orthogonal basis clustering and local manifold distribution constraints.

[0036] Furthermore, the processing procedure of the closed-loop optimization module includes: To establish an overall optimization objective function that integrates local confidence constraints, tensor low-rank constraints, and orthogonal basis clustering, the overall objective function is as follows: ; Subject to the following constraints (st): ; Among them, the first five items The constraints respectively ensure that the available similarity graph, the complete similarity graph, and the self-representation reconstruction equation satisfy the probabilistic simplex distribution; This ensures the normalization of the fusion weights of each view; and This ensures the independence and mutual exclusion of the orthogonal basis space and clustering indicators; The decoupling decomposition path of multi-view self-representation tensors is limited.

[0037] Furthermore, it also includes: Alternating variable optimization updates are employed, introducing auxiliary variables and using the Enhanced Lagrange Multiplier Method (ALM) to transform the overall problem into multiple subproblems for alternating updates: 1. Update and complete the matrix Similarity map available Updates are calculated directly based on analytical solutions.

[0038] 2. Update the complete data matrix With low-rank essential tensor proxy variables : For the introduced complete data matrix Closed-form updates are performed by solving the standard Sylvester equations. Proxy variables for low-rank essential tensors Tensor singular value decomposition (t-SVD) is performed on the stacked tensors in the Fourier transform domain, and the weighted singular value threshold shrinkage operator is applied to reconstruct the updated low-rank tensor.

[0039] 3. Update the self-representation error matrix With noise tensor : Targeting the The error and noise terms of the norm constraint are calculated using the per-row shrinkage operator. -norm thresholding) is used for sparsification updates, automatically filtering out structural noise caused by outliers.

[0040] 4. Update the local structure Laplace matrix: based on the highly reliable consensus graph generated by the fusion in the current iteration step. Calculate the degree matrix , and according to Dynamically update the globally consistent Laplacian matrix.

[0041] 5. Update Orthogonal basis moments With pseudo-clustering label matrix For those subject to orthogonal constraints ( , The subproblem of ) is solved by constructing an intermediate target matrix and performing singular value decomposition (SVD) on it. The left and right singular matrices obtained by the decomposition are multiplied to satisfy the variable update in the normal coherent space.

[0042] 6. Update the complete similarity map Under the combined effect of self-representation error constraints and confidence map Hadamard product filtering, the following is achieved: We perform row-by-row optimization and ensure that the probabilistic simplex constraint (non-negative elements and row sum of 1) is satisfied through the projection operator.

[0043] 7. Update Lagrange multipliers and penalty parameters: At the end of each iteration, increment the penalty parameter according to a fixed rule. This accelerates algorithm convergence.

[0044] Determine if the algorithm has converged (e.g., residuals are less than a threshold or the maximum number of iterations has been reached). After convergence, determine the convergence based on the feature weights. All features are sorted in descending order, and the most important features are selected. Finally, the K-means clustering algorithm is used to cluster the selected features, and the final clustering results are output.

[0045] In this embodiment, after extracting the dimensionality-reduced and pure features, in addition to using the K-means algorithm, other algorithms such as spectral clustering or hierarchical clustering can also be used to output the final result.

[0046] This embodiment also provides an incomplete multi-view clustering method based on confidence graphs and tensor feature selection, including: S1. Obtain incomplete multi-view data and construct an initial similarity map and confidence map based on the available data in each view; S2. Construct the first optimization subproblem containing self-representation reconstruction terms and confidence graph filtering terms, introduce a data completion matrix to fill in missing instances, and iteratively update the complete similarity graph of each view; S3. Stack the updated complete similarity maps of each view into a self-representation tensor, construct a second optimization subproblem containing tensor low-rank constraints, decompose the self-representation tensor into an essential tensor and a noise tensor, and apply a weighted tensor nuclear norm constraint to the essential tensor to capture the high-order correlation between views. S4. Construct a global Laplacian matrix based on the essential tensor, and construct a third optimization subproblem that includes feature projection, orthogonal basis clustering and local manifold preservation, and solve for the feature projection matrix; S5. Integrate the first optimization sub-problem, the second optimization sub-problem, and the third optimization sub-problem into a general objective function, perform joint optimization using the alternating direction multiplier method, and feed back the solution result of the feature projection matrix to S2 during the iteration process to dynamically adjust the update direction of the similarity map.

[0047] S6. After the iteration converges, the original features are sorted according to the row vector norm of the feature projection matrix, a subset of discriminative features is selected, and the clustering algorithm is input to output the final clustering result.

[0048] Specifically, including: Step 1: Extracting multi-view data and initializing local graphs.

[0049] For incomplete multi-view data containing missing instances, the available data matrix for each view is extracted. Based on the extracted available data, an initial similarity map between available samples is calculated using Euclidean distance; simultaneously, based on the available data... A binary indicator matrix is ​​constructed based on nearest neighbor relationships. A confidence graph is generated by calculating the number of common neighbors and standardizing it by percentage. The confidence graph is used to quantify the reliability of local connections between samples.

[0050] Step 2: Dynamically reconstruct the complete similarity map by combining confidence filtering.

[0051] Based on the self-representation learning mechanism and the missing data imputation matrix, a self-representation reconstruction model for incomplete data is established to dynamically learn the complete global similarity map. During the reconstruction process, the confidence map generated in step one is used as a structure filter, and the Hadamard product (element-wise multiplication) is used to perform a masking operation on the reconstructed similarity matrix to filter out discrete points and pseudo-topological connections caused by high-dimensional noise.

[0052] Step 3: Tensor decoupling and extraction of global high-order consistency relationships.

[0053] The reconstructed similarity maps of each view are stacked and rotated along the view dimension to construct a self-representation tensor. The self-representation tensor is decoupled and decomposed into an essential tensor and a noise tensor, and a weighted tensor kernel norm (WTNN) is applied to the essential tensor. Through tensor singular value decomposition (t-SVD) and adaptive singular value weighting, local high-order noise is suppressed, and the global high-order consistent manifold structure across views is extracted.

[0054] Step 4: Unsupervised feature selection based on orthogonal basis clustering and preservation of local structure.

[0055] A globally consistent Laplacian matrix is ​​generated using the essential tensor. Guided by both the globally consistent structure and the local manifold distribution of the data, a feature selection optimization objective is established. The recovered complete data matrix is ​​projected onto a low-dimensional cluster space through orthogonal basis decomposition, and then... Norm constraints induce row sparsity in the feature projection matrix, thereby filtering out redundant features.

[0056] Step 5: Joint closed-loop optimization and final clustering output.

[0057] Local confidence constraints, tensor low-rank constraints, and orthogonal basis clustering are integrated into a strictly constrained overall optimization objective function, which is then solved iteratively using the Alternating Direction Multiplier Method (ADMM). Purified features extracted through feature selection are fed back into the similarity map construction process, forming a closed-loop optimization. After convergence, the features are sorted in descending order based on their weights in the feature projection matrix, and a subset of discriminative features is selected and input into the K-means algorithm, outputting the final clustering results.

[0058] To more clearly illustrate the technical solution of the present invention, specific embodiments are provided below for description: Figure 1 The overall architecture and data flow diagram of the incomplete multi-view clustering method based on confidence graphs and tensor feature selection include the following key technical steps: 1. Data Preprocessing and Local Graph Initialization: First, for incomplete multi-view data containing missing instances, ( Figure 1 The dashed box at the top indicates missing data. In this embodiment, the Notting-Hill dataset is used to indicate three specific feature views: 2000-dimensional intensity features, 3304-dimensional Local Binary Pattern (LBP) features, and 6750-dimensional Gabor features. The usable data matrix for each view is extracted at a 30% missing data rate. Since constructing a map directly on incomplete data introduces significant errors, this embodiment utilizes only the extracted usable data matrix. Constructing a basic available affinity graph Simultaneously, a confidence graph is constructed based on the nearest neighbor relationships of available data. The confidence plot quantifies the reliability of inter-sample connections in the form of percentage shares, serving as structural prior information for subsequent steps at this stage.

[0059] 2. Dynamic Reconstruction and Structural Filtering of Similarity Graphs: Utilizing the initialized available similarity graph and confidence graph, the graph is expanded to a complete graph. This involves reconstructing the complete similarity graph for each view. When, the confidence graph generated in step 1 is used. As a "structural filter," specifically, it uses the Hadamard product (element-wise multiplication) of matrices and the confidence graph to suppress outliers and high-dimensional noise, preventing the propagation of local pseudo-topological connections during the missing data recovery process, thereby obtaining a high-purity, complete similarity matrix. .

[0060] 3. Tensor quantization and global high-order relation extraction: In order to capture high-order correlations across views, the complete similarity matrix of all views is obtained. Stacking along the view dimension and rotating by 90 degrees constructs a third-order self-representation tensor. ).

[0061] Since slight noise from interpolation padding may still remain in the complete graph, this embodiment further decouples and decomposes the self-representation tensor into an "essential tensor". ) and a "noise tensor" Perform tensor singular value decomposition (t-SVD) based on weighted tensor nuclear norm (WTNN) on the essential tensor, thereby distilling a higher-order consistent manifold structure across the global tensor space.

[0062] 4. Joint Optimization, Data Recovery, and Feature Selection: The decomposed essential tensor is then used to generate the Laplacian matrix, guiding the feature selection and clustering modules. Simultaneously, orthogonal-constrained completion and alignment matrices are introduced, combined with the original incomplete multi-view data. Dynamically recover complete sample feature information (i.e. Guided by both preserving local manifold structure and global tensor consistency, orthogonal basis decomposition is performed on the recovered complete data to solve for the feature projection matrix, thereby filtering out redundant and noisy features and completing unsupervised feature selection.

[0063] 5. Closed-loop feedback and final clustering output: This embodiment is not a unidirectional pipeline, but a dynamic feedback system. The purified features extracted through unsupervised feature selection are fed back into the system (e.g., ...). Figure 1 The feedback loop (the one indicated by the rightmost downward arrow) iteratively updates and optimizes the construction of the similarity map and confidence map. Once the entire alternating optimization model converges, the features are sorted in descending order based on the weights of the final feature projection matrix. The most discriminative feature subset is then selected and input into the K-means algorithm, ultimately outputting accurate clustering results.

[0064] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An incomplete multi-view clustering system based on confidence graph and tensor feature selection, characterized in that, include: The data preprocessing and local graph initialization module is used to extract the available data matrix under each view for incomplete multi-view data containing missing instances, and to construct an initial similarity map, confidence map and complete similarity map to be learned based on the available data matrix. The confidence-guided graph dynamic reconstruction module is used to iteratively update and reconstruct the complete similarity graph based on a self-representation learning mechanism, combined with the confidence graph as a structural filter. The tensor decoupling and higher-order relation extraction module is used to stack the complete similarity graphs after each view reconstruction to construct a self-representation tensor and extract the global higher-order consistent manifold structure across views. The feature selection and clustering module is used to generate a globally consistent Laplacian matrix based on the global high-order consistent manifold structure, and to establish a feature selection optimization objective by combining the local manifold distribution of the data. It also filters discriminative feature subsets through orthogonal basis decomposition and sparse constraints, and outputs clustering results. The closed-loop optimization module is used to integrate the optimization objectives of the confidence-guided graph dynamic reconstruction module, tensor decoupling and higher-order relation extraction module, and feature selection and clustering module into a general objective function, and solve iteratively.

2. The incomplete multi-view clustering system based on confidence graph and tensor feature selection according to claim 1, characterized in that, Constructing the confidence graph includes: Based on the available data matrix, a binary indicator matrix is ​​constructed using K-nearest neighbor relationships, and the inner product of the binary indicator matrix is ​​calculated to obtain a matrix of common neighbors. The common neighbor number matrix is ​​percentage-normalized to generate the confidence graph, which is used to quantify the reliability of local connections between samples.

3. The incomplete multi-view clustering system based on confidence graph and tensor feature selection according to claim 1, characterized in that, The confidence-guided graph dynamic reconstruction module utilizes the confidence graph as a structure filter in the following ways: In the objective function of reconstructing the complete similarity map, the Hadamard product operator is introduced to perform element-wise multiplication between the complete similarity map to be learned and the aligned confidence map to form a mask operation, which forcibly reduces the weights of the connections that are judged as low reliability by the confidence map.

4. The incomplete multi-view clustering system based on confidence graph and tensor feature selection according to claim 1, characterized in that, Extracting the global higher-order consistent manifold structure across the views includes: The self-representation tensor is decoupled and decomposed into an essential tensor and a noise tensor. A weighted tensor kernel norm constraint is applied to the essential tensor, and the global high-order consistent manifold structure across the view is extracted.

5. The incomplete multi-view clustering system based on confidence graph and tensor feature selection according to claim 4, characterized in that, Applying a weighted tensor nuclear norm constraint to the essential tensor is specifically as follows: ; In the formula, It is an essential tensor; Represents the number of views; This represents the total number of samples actually processed. For the essential tensor in the Fourier transform domain, the first... One slice; The essential tensor in the Fourier transform domain is the first... The first slice Large singular values; These are the weighting coefficients; Let represent the weighted tensor nuclear norm.

6. The incomplete multi-view clustering system based on confidence graph and tensor feature selection according to claim 1, characterized in that, The feature selection optimization objective established by the feature selection and clustering module is: ; In the formula, For the complete data matrix; This refers to the feature projection matrix or the feature weight matrix; It is an orthogonal basis matrix; This is a pseudo-clustering label matrix; for Norm constraints; It is a globally consistent Laplace matrix; It is the Frobenius norm; , All are balance parameters; Let be the trace of the matrix.

7. An incomplete multi-view clustering method based on confidence graph and tensor feature selection, applied to the incomplete multi-view clustering system based on confidence graph and tensor feature selection as described in any one of claims 1-6, characterized in that, include: S1. Obtain incomplete multi-view data and construct an initial similarity map and confidence map based on the available data in each view; S2. Construct the first optimization subproblem containing self-representation reconstruction terms and confidence graph filtering terms, introduce a data completion matrix to fill in missing instances, and iteratively update the complete similarity graph of each view; S3. Stack the updated complete similarity maps of each view into a self-representation tensor, construct a second optimization subproblem containing tensor low-rank constraints, decompose the self-representation tensor into an essential tensor and a noise tensor, and apply a weighted tensor nuclear norm constraint to the essential tensor to capture the high-order correlation between views. S4. Construct a global Laplacian matrix based on the essential tensor, and construct a third optimization subproblem that includes feature projection, orthogonal basis clustering and local manifold preservation, and solve for the feature projection matrix; S5. Integrate the first optimization sub-problem, the second optimization sub-problem, and the third optimization sub-problem into a general objective function, perform joint optimization using the alternating direction multiplier method, and feed back the solution result of the feature projection matrix to S2 during the iteration process to dynamically adjust the update direction of the similarity map.

8. The incomplete multi-view clustering method based on confidence graph and tensor feature selection according to claim 7, characterized in that, Also includes: After the iteration converges, the original features are sorted according to the row vector norm of the feature projection matrix, a subset of discriminative features is selected, and the clustering algorithm is input to output the final clustering result.