Multi-view subspace clustering method for enhancing tensor Schatten P-norm

By enhancing the multi-view subspace clustering method of tensor Schatten P-norm, the problems of noise interference and model complexity in multi-view subspace clustering are solved, and more efficient data analysis and clustering performance are achieved.

CN120687857APending Publication Date: 2025-09-23ANHUI POLYTECHNIC UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510770461.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing multi-view subspace clustering methods are seriously affected by noise interference and outliers when processing complex multi-source heterogeneous data. In addition, unreasonable parameter settings lead to complex model solving process and ignore the important structural information of singular values.

Method used

A multi-view subspace clustering method with enhanced tensor Schatten P-norm is adopted. By constructing an initial denoising model and applying weighted tensor Schatten P-norm, the model is optimized with alternating direction multiplier method to remove noise, assign weights to different singular values, and integrate the high-order correlation of multi-view data.

Benefits of technology

It effectively removes data noise, reduces the complexity of model solution, highlights data structure characteristics, improves clustering performance, and demonstrates excellent clustering effects on a variety of data sets through experimental verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687857A_ABST
    Figure CN120687857A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view subspace clustering method for enhancing tensor Schatten P-norm, and relates to the field of multi-view data subspace clustering, the method comprises the following steps: firstly, modeling noise data, and simultaneously considering Laplace noise (l1-norm represents loss) and Gaussian noise (l2, 1-norm represents loss) in the data; noise pollution in the data is effectively removed, and a cleaner tensor structure is obtained. Then, a weighted tensor Schatten P-norm is applied to the denoised tensor, which not only approaches the original structure feature of the tensor, but also fully considers the contribution of different singular values to the clustering performance. According to the method, a tensor denoising model and a weighted tensor Schatten P-norm are integrated into a unified framework, so that effective analysis of complex data is realized, different singular values are endowed with different weights, and remarkable structural features in the data are highlighted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multi-view data subspace clustering, and in particular, relates to a multi-view subspace clustering method based on an enhanced tensor Schatten P-norm. Background Art

[0002] Subspace clustering is widely used due to its excellent clustering capabilities. Within the subspace clustering framework, sparse subspace clustering (SSC) and low-rank representation (LRR) are representative algorithms that primarily reveal the underlying subspace structure through single-view data modeling. However, single-view learning struggles with multi-source heterogeneous data scenarios, and its clustering performance is susceptible to multiple influences, including noise, outliers, and complex data distributions.

[0003] To address the limitations of single-view learning in complex data scenarios, multi-view subspace clustering (MSC) has become a research hotspot by fusing local structural information across views. Its core is to integrate the structural features of multi-source heterogeneous data and improve clustering performance through collaborative representation learning. For example, existing technologies have proposed a convex representation MSC method that ensures the consistency of view information by constructing a consensus matrix. Another approach is to combine sparse and low-rank constraints to explore the effective information of views through multiple regularization constraint strategies.

[0004] As research deepens, the MSC method has gradually evolved from traditional matrix modeling to a tensor analysis paradigm to more fully capture high-order coupling relationships between views. For example, existing techniques treat multi-view coefficient matrices as horizontal slices of a tensor and use a tensor nuclear norm (TNN) minimization strategy to effectively extract high-order information between views. Furthermore, a MSC method based on tensor singular value decomposition has been proposed. By concatenating multi-view coefficient matrices into a tensor and applying TNN constraints, it achieves joint learning of high-order features between views and combines spectral clustering to obtain clustering results.

[0005] However, most of the above MSC methods use 2,1 -norm to constrain noise information. This assumption is only applicable to specific noise distributions and cannot accurately characterize the complex and diverse noise characteristics in real scenarios. In addition, the existing MSC method has unreasonable parameter settings for the constructed denoising model and does not adjust them, which makes the model solution process complicated. In addition, existing research shows that larger singular values ​​often contain the most important structural information in the data, while smaller singular values ​​are usually composed of noise data. Treating all singular value information equally will ignore the impact of matrix prior information on clustering performance. Summary of the Invention

[0006] In response to the problems in the related art, the present invention proposes a multi-view subspace clustering method based on an enhanced tensor Schatten P-norm to overcome the above technical problems existing in the existing related art.

[0007] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0008] The present invention is an enhanced tensor Schatten P-norm for multi-view subspace clustering (ETSPMSC) method, comprising the following steps:

[0009] S1. Constructing an initial ETSPMSC model; specifically comprising the following steps:

[0010] S11, setting multi-view data; constructing a similarity matrix of the view data set to be processed; and then constructing a probability transfer matrix based on the similarity matrix;

[0011] S12, constructing an initial denoising model in conjunction with the probability transfer matrix; the initial denoising model includes Laplace noise (l1-norm represents loss) and Gaussian noise (l 2,1 -norm represents loss);

[0012] S13, applying the weighted tensor Schatten P-norm to the initial denoising model to obtain an initial ETSPMSC model;

[0013] S2. updating and adjusting the variables in the initial ETSPMSC model using the alternating direction multiplier method (ADMM) in conjunction with the multi-view data to optimize the initial ETSPMSC model to obtain an optimized ETSPMSC model;

[0014] S3. Using the optimized ETSPMSC model to perform denoising on the view dataset to be processed, to obtain a denoised view tensor;

[0015] S4. Stack the denoised view tensors into a matrix and perform clustering on the matrix using spectral clustering.

[0016] Preferably, the S12 includes the following steps:

[0017] S121, decomposing the probability transfer matrix into a probability matrix and an error matrix;

[0018] S122, applying a low-rank constraint to the probability matrix to obtain a low-rank constrained probability matrix;

[0019] S123, stack the probability matrices in S121 along the direction of modulo 3 to generate a probability matrix tensor, and explore various noise pollutions from a high-dimensional level;

[0020] S124. Considering different noise information in the data at the same time, the error matrix in S121 is deleted, and an initial denoising model is constructed to characterize and remove noise pollution.

[0021] Preferably, the low-rank constraint formula is as follows:

[0022]

[0023] Where: E v represents the error matrix, P v represents the probability transfer matrix, Z v represents the probability matrix, ||·|| * represents the low-rank constraint, ||·|| 2,1 Indicates l 2,1 -norm constraint, V represents the total number of views, and λ is a smoothness term.

[0024] Preferably, the initial denoising model formula in S123 is as follows:

[0025]

[0026] Where: E represents the Gaussian noise tensor, N represents the Laplace noise tensor; P is the probability transfer tensor; Z is the probability tensor, which is formed by stacking the probability matrix along the modulo 3 direction; λ1 and λ2 are regularization parameters used to constrain E and N, ||·|| TNN is the tensor core norm constraint, and ||·||1 is the l1-norm constraint of the tensor.

[0027] Preferably, the specific process of applying the weighted tensor Schatten P-norm to the initial denoising model in S13 is as follows:

[0028] The weighted tensor Schatten P-norm is used to replace the TNN in the initial denoising model, and graph regularization constraints are imposed on the view to ensure the local structural characteristics of the view.

[0029]

[0030] Where tr(·) represents the trace of the matrix, L v is the Laplace matrix, represents the weighted tensor Schatten P-norm, and λ3 is the regularization parameter.

[0031] Preferably, said S2 comprises the following steps:

[0032] S21, setting an auxiliary tensor; using the auxiliary tensor to perform equivalent transformation on the initial ETSPMSC model to obtain a transformed ETSPMSC model;

[0033] S22, solving the transformed ETSPMSC model using an alternating direction multiplier method to obtain a solved ETSPMSC model;

[0034] S23. Using an alternating direction multiplier method to update the variables in the solved ETSPMSC model.

[0035] Preferably, the S23 includes the following steps:

[0036] S231, setting adjustable parameters; inputting the multi-view data and the adjustable parameters into the solved ETSPMSC model;

[0037] S232, initializing the parameters in the solved ETSPMSC model to obtain an initialized ETSPMSC model;

[0038] S233. Set a convergence formula; use the alternating direction multiplier method to calculate the solution formula of multiple variables in the initialized ETSPMSC model; then update the multiple variables in the initialized ETSPMSC model in conjunction with the solution formula of the multiple variables until the convergence formula is established, and obtain a probability matrix.

[0039] Preferably, the multiple variables in the initialized ETSPMSC model in S233 include a probability tensor, an auxiliary tensor, a Gaussian noise tensor, a Laplace noise tensor, a Lagrange multiplier, and a penalty parameter.

[0040] Preferably, the convergence formula is as follows,

[0041] ||PZEN|| ∞ <ε, ||ZT|| ∞ <ε (8)

[0042] Among them, ||PZEN|| ∞ <ε and ||ZT|| ∞ <ε indicates that the infinite norm of the residual should be less than the threshold ε; this indicates that the decomposition result is close enough to the original data and the error is within the controllable range; ε=10 -7 , when the convergence condition is less than ε, the algorithm is considered to have converged; T represents the auxiliary tensor; ||·|| ∞ is the infinity norm of the tensor.

[0043] A multi-view subspace clustering system based on enhanced tensor Schatten P-norm includes a denoising model building module, an ETSPMSC model building module, and an ETSPMSC model optimization module.

[0044] The present invention has the following beneficial effects:

[0045] 1. In the process of modeling noise data in the present invention, the Laplace noise (l1-norm represents loss) and Gaussian noise (l 2,1 -norm representation loss), effectively removing noise pollution in the data and obtaining a cleaner tensor structure; and iteratively adjusting and optimizing the variables (parameters) of the constructed denoising model, greatly reducing the solution complexity of the denoising model; then, applying the weighted tensor Schatten P-norm to the denoised tensor, not only is it closer to the original information of the tensor, but also fully considers the contribution of singular values ​​to the clustering performance; by integrating the tensor denoising model and the weighted tensor Schatten P-norm into a unified framework, not only can effective analysis of complex data be achieved, but also different weights are assigned to different singular values, thereby highlighting the significant structural features in the data; by introducing Markov chain representation learning, the computational complexity of the algorithm is effectively reduced; in addition, a comprehensive experimental evaluation is carried out on five datasets, including performance comparison, parameter analysis, convergence analysis and running time comparison. The experimental results fully demonstrate the excellent clustering performance of the ETSPMSC model.

[0046] 2. In the present invention, a third-order tensor of the transition probability matrix is ​​constructed and rotated to explore the high-order correlation in multi-view data; on this basis, the tensor is singular value weighted and rank function approximated to effectively learn the prior knowledge of the view.

[0047] 3. In the present invention, an optimization strategy based on the Alternating Direction Method of Multipliers is designed to ensure efficient solution of the objective function.

[0048] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1Schematic diagram of the ETSPMSC framework in the present invention, wherein: Construct Transition Probability Matrix: construct probability transition matrix; Dataset: dataset; Similarity Matrix: similarity matrix; Transition Probability Matrix: probability transition matrix; Rotate: tensor rotation; t-SVD based weighted Schatten P-norm: weighted tensor Schatten P-norm based on tensor singular value decomposition; GraphConstraint: graph regularization;

[0051] Figure 2 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 0.001 in the present invention;

[0052] Figure 3 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 0.005 in the present invention;

[0053] Figure 4 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 0.01 in the present invention;

[0054] Figure 5 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 0.03 in the present invention;

[0055] Figure 6 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 0.05 in the present invention;

[0056] Figure 7 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 0.1 in the present invention;

[0057] Figure 8 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3=1 in the present invention;

[0058] Figure 9 Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 10 in the present invention;

[0059] Figure 10Schematic diagram of the distribution of the effects of parameters λ1, λ2, and λ3 on ACC on the COIL-20 dataset when λ3 = 100 in the present invention;

[0060] Figure 11 Schematic diagram of the influence of the weight parameter ω on the clustering performance on BBCsport in the present invention;

[0061] Figure 12 Schematic diagram of the influence of the weight parameter ω on the clustering performance on BBC4view in the present invention;

[0062] Figure 13 Schematic diagram of the effect of the weight parameter ω on the clustering performance on COIL-20 in the present invention;

[0063] Figure 14 Schematic diagram of the effect of parameter p on ACC and NMI on BBCsport in the present invention;

[0064] Figure 15 Schematic diagram of the effect of parameter p on ACC and NMI on BBC4view in the present invention;

[0065] Figure 16 Schematic diagram of the effect of parameter p on ACC and NMI in UCI in the present invention;

[0066] Figure 17 Schematic diagram of the convergence curve of the ETSPMSC algorithm in the present invention on the COIL-20 benchmark data set; where Error: error rate; Iteration: number of iterations;

[0067] Figure 18 Schematic diagram of the convergence curve of the ETSPMSC algorithm in the present invention on the UCI benchmark dataset;

[0068] Figure 19 Schematic diagram of the convergence curve of the ETSPMSC algorithm in the present invention on the Prokaryotic benchmark dataset. DETAILED DESCRIPTION

[0069] The following will clearly and completely describe the technical solutions in the embodiments of the invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0070] Example 1

[0071] See also Figure 1This embodiment is a multi-view subspace clustering method for enhancing the tensor Schatten P-norm, comprising the following steps:

[0072] 1.1 ETSPMSC model construction

[0073] In the Markov chain-based MSC framework, we first use the original view matrix X V Construct similarity matrix S v , used to calculate the similarity between different data points in the original view; σ is the standard deviation, Represents the i-th row data of the original view matrix, Represents the j-th row data of the original view matrix; then through Construct probability transfer matrix P v , where D v It is a diagonal matrix, and the elements on the diagonal are obtained by adding each row of the similarity matrix, that is, Indicates D v The data in row i and column i; Indicates S v In order to process the redundant information in the view, the probability transfer matrix P v Decomposed into two parts: probability matrix Z v And the error matrix E v At the same time, Z v The low-rank constraint is imposed, and its formula is as follows:

[0074]

[0075] Where V represents the total number of views, ||·|| * represents the low-rank constraint, ||·|| 2,1 Indicates l 2,1 -norm constraint, λ is a smoothing term to avoid the error matrix E v Excessive loss; E v represents the error matrix, P v represents the probability transfer matrix, Z v represents the probability matrix.

[0076] On this basis, the shared matrix in Equation (9) is stacked along the modulo 3 direction to generate a 3rd-order probability tensor to explore the high-order correlation between views.

[0077] In addition, Equation (9) cannot fully remove the noise information in the data. To this end, a denoising model is constructed to better represent the noise information. Unlike the tensor low-rank representation method that only considers Laplace noise, it considers both Laplace and Gaussian noise to more accurately reveal the subspace structure of the sample tensor:

[0078]

[0079] Where E represents the Gaussian noise tensor, N represents the sparse noise tensor (Laplace noise tensor), Z represents the probability tensor, P is the probability transfer tensor, λ1 and λ2 are regularization parameters used to constrain E and N, ||·|| TNN is the tensor nuclear norm constraint, ||·|| 2,1 is the tensor l 2,1 -norm constraint, ||·||1 is the l1-norm constraint of the tensor.

[0080] In addition, using the weighted tensor SchattenP-norm instead of TNN can better learn the prior knowledge of the view and maintain the overall structure of the probability matrix through manifold learning:

[0081]

[0082] Where tr(·) represents the trace of the matrix, L v is the Laplace matrix, represents the weighted tensor Schatten P-norm, and λ3 is the regularization parameter.

[0083] 1.2 Optimization of ETSPMSC

[0084] 1.2.1 To simplify the solution process of Z, an auxiliary tensor T is introduced. Equation (11) can be equivalently transformed into the following form:

[0085]

[0086] stP=Z+E+N,T=Z (12)

[0087] Since Equation (12) is non-convex, it can be solved by the alternating direction multiplier method. The specific formula is:

[0088]

[0089] Where μ is the penalty parameter, C and L are Lagrange multipliers, is the F-norm constraint, and <·,·> represents the inner product of the matrices.

[0090] 1.2.2 Use the alternating direction multiplication method to update Z, T, E and N, that is, fix other irrelevant variables and update single variables in turn.

[0091] 1.2.2.1 Update Z

[0092] After fixing other variables, Z can be expressed as:

[0093]

[0094] Formula (14) only selects the items related to the variable Z, and by minimizing the parameters related to Z, the local optimal solution of Z (Z k+1 ).

[0095] To Z v After taking the partial derivative and setting it to 0, we can get the closed-form solution:

[0096]

[0097] in, and I represents the identity matrix; Λ v represents the Lagrange multiplier.

[0098] 1.2.2.2 Update T

[0099] After fixing other variables, T can be expressed as:

[0100]

[0101] Formula (16) only selects the items related to the variable T, and obtains the local optimal solution of T (T k+1 ).

[0102] The solution is:

[0103]

[0104] Among them, Γ is the threshold operation parameter, which can be used to reduce the dimension of the tensor. is the threshold parameter.

[0105] 1.2.2.3 Update E

[0106] After fixing other variables, E can be expressed as:

[0107]

[0108] The closed-form solution is:

[0109]

[0110] Wherein, formula (19) is a piecewise function, which is used to process each column of matrix D separately. :,i Represents the data in the i-th column of matrix D to determine the updated value of E

[0111] 1.2.2.4 Update N

[0112] After fixing other variables, N can be expressed as:

[0113]

[0114] make The solution of Equation (12) can be obtained by using the element-wise contraction operator ∑ n (·) to obtain:

[0115]

[0116] 1.2.2.5 Update C, L and μ:

[0117]

[0118] Where ρ = 2 to accelerate convergence; μ max Indicates the maximum value of the penalty parameter.

[0119] Algorithm 1 summarizes the complete optimization process of ETSPMSC as follows,

[0120]

[0121] where ε = 10 -7 , when the convergence condition is less than ε, the algorithm is considered to have converged; ||·|| ∞ is the infinity norm of the tensor.

[0122] 1.3 Algorithm Complexity Analysis

[0123] For the ETSPMSC algorithm, its algorithm complexity is mainly concentrated on the update of the four sub-problems Z, E, N and T. For Z, the complexity is mainly concentrated on the matrix inversion, and its complexity for all views is O(VN 3 For N, the complexity of acting on all views is O(VN 2 ). E has a closed-form solution with complexity O(VN 2 For T, a weighted tensor Schatten P-norm minimization strategy is required, which has a complexity of O(VN 2 log(VN)+V 2 N 2 Considering the number of iterations K and the total condition of 1#V log(N)N, the total time complexity is O(KVN 2 (N+log(VN)).

[0124] 2 Experiments

[0125] In this experiment, five multi-view datasets are selected to compare ETSPMSC with some state-of-the-art clustering algorithms, and the results are analyzed using six commonly used metrics, namely, accuracy (ACC), normalized mutual information (NMI), adjusted Rand index (AR), F-score, precision, and recall.

[0126] 2.1 Experimental Dataset

[0127] To verify the performance of ETSPMSC, five different datasets were selected in this technical proposal. The following is a brief introduction to these five datasets:

[0128] BBCsport and BBC4view datasets: Both datasets are collected from the sports news section of the BBC News website. The BBC4view dataset contains 685 sports news documents, and four different types of features are extracted from each document; the BBCsport dataset contains 544 sports news documents, and two types of feature representations are extracted from each document.

[0129] COIL-20: This dataset consists of 20 categories, each containing 72 images, for a total of 1,440 grayscale images. Three different types of features (intensity, LBP, and Gabor) were extracted.

[0130] UCI: This dataset contains 2,000 handwritten digit samples from 10 categories, ranging from 0 to 9. The experiment uses three feature representation methods: pixel average feature, Fourier coefficient feature, and data morphological feature.

[0131] Prokaryotic: This dataset contains 551 prokaryotic samples. Three types of features were extracted for characterization: text features, gene report features, and proteome comparison features.

[0132] Table 1 briefly summarizes the contents of the above datasets.

[0133] Table 1 Dataset introduction

[0134] Table 1Introduction to the datasets

[0135]

[0136] 2.2 Comparison of Algorithms

[0137] To comprehensively evaluate the performance of the ETSPMSC algorithm, this technical solution selected ten classic clustering algorithms as comparison benchmarks. The following is a brief introduction to the comparison algorithms:

[0138] (1) SSC: Clustering single-view data through sparse constraints to effectively capture the sparse structure of the data.

[0139] (2) LRR: Use low-rank representation to restore the subspace structure of the original data, thereby better revealing the intrinsic low-rank characteristics of the data.

[0140] (3) WTNNM: Introduces the weighted tensor nuclear norm to explore high-order correlations between views and enhance the representation ability of multi-view data.

[0141] (4) ETLMSC: Combining probabilistic transfer tensors and sparsity constraints, it deeply mines view information and improves multi-view clustering performance.

[0142] (5) LT-MSC: Enhance and improve the traditional MSC by using low-rank tensor constraints to further improve the clustering effect.

[0143] (6) t-SVD-MSC: Combines tensor singular value decomposition to learn high-order tensor subspaces and effectively process high-dimensional multi-view data.

[0144] (7) HLR-MVS: It uses a hypergraph structure to capture the correlation between views and the local manifold structure, providing richer structural information for multi-view clustering.

[0145] (8) RMSC: The Markov chain method is used to optimize the probability transfer matrix to better model the probabilistic relationship between data and improve clustering performance.

[0146] (9)LSGMC: It uses symmetry constraints and matrix decomposition methods to explore the consistency between views and further optimize the multi-view clustering results.

[0147] (10) IiAGLMC: Constructs view structure based on subspace learning method and introduces inclusiveness-induced regularizer to reconstruct graph structure to improve the robustness of multi-view clustering.

[0148] 2.3 Clustering Performance Comparison

[0149] This section selects the optimal parameters for all compared algorithms through parameter tuning. Table 2 shows the clustering performance comparison results on five benchmark datasets, with the best performance indicated in bold and the suboptimal performance indicated by underline. The performance comparison of ETSPMSC with the other ten algorithms is as follows:

[0150] 1) ETSPMSC shows significant performance advantages on all datasets. Specifically: it achieves zero-error clustering on the BBCsport dataset. On the BBC4view dataset, the ACC, NMI, and F-Score indicators are improved by 2.7%, 5.6%, and 5.0% respectively compared with the second-best method WTNNM. On the COIL-20 dataset, the clustering index is improved by about 3% on average compared with LSGMC. On the UCI dataset, the tensor learning-based methods perform well overall. On the Prokaryotic dataset, the clustering index is improved by 4.9%, 22.3%, and 9.5% respectively compared with RMSC, which is also a Markov chain method.

[0151] 2) ETSPMSC outperforms single-view methods such as SSC and LRR, thanks to the fact that the MSC algorithm can more comprehensively utilize the correlation properties between different views.

[0152] 3) Compared with Markov chain methods such as RMSC and ELTMSC, ETSPMSC effectively reduces complex noise interference and maintains local structural characteristics through dual noise modeling (Laplacian + Gaussian) and manifold regularization strategy.

[0153] 4) LSGMC and liAGLMC are matrix-based methods and cannot fully explore high-order correlations between views.

[0154] 5) The performance advantage of ETSPMSC over other tensor methods mainly comes from its full consideration of the significant differences between singular values ​​and the precise approximation of the original information by the coefficient matrix achieved through the Schatten P-norm constraint.

[0155] Table 2 Comparison of algorithm clustering performance

[0156] Table 2Comparison of clustering performance

[0157]

[0158]

[0159] 2.4 Parameter Analysis

[0160] The ETSPMSC method involves optimizing five key parameters: the regularization parameters λ1, λ2, and λ3, the Schatten P-norm parameter p, and the weight coefficient ω. The tuning ranges for each parameter are as follows: the regularization parameters λ1, λ2, and λ3 are selected from the discrete value set [0.001, 0.005, 0.01, 0.03, 0.05, 0.1, 1, 10, 100]; the Schatten P-norm parameter p is selected from the interval [0.1, 0.2, ..., 1] in steps of 0.1; and the weight coefficient ω is selected from the interval [0, 100].

[0161] Figure 2-Figure 10 This is the parameter sensitivity analysis of parameters λ1, λ2 and λ3. It can be seen from the figure that ETSPMSC shows stability in a large range of parameter changes. In addition, Figure 11 、 Figure 12 、 Figure 13 The effect of the choice of singular value weights on performance is shown, from which it can be seen that clustering performance tends to be better when larger singular values ​​are scaled less. Figure 14 、 Figure 15 、 Figure 16 The effect of different p on clustering performance is shown. By scaling with different p values, the results can be made closer to the original rank function.

[0162] For the BBCsport, BBC4view, COIL-20, UCI, and Prokaryotic datasets, the optimal parameters of λ1, λ2, λ3, and p for each dataset are as follows: (10, 0.01, 10, 0.7), (0.1, 0.001, 1, 0.6), (0.001, 0.005, 0.1, 0.6), (0.03, 0.001, 100, 0.9), and (0.03, 0.001, 100, 1). The optimal parameter ω for each dataset is: [1, 10], [0.1, 1, 10, 100], [0.1, 100, 0.1], [1, 1, 1], and [1, 0.1, 10].

[0163] 2.5 Time Consumption Analysis

[0164] In the experiments, eight multi-view clustering algorithms were compared with ETSPMSC, and their runtimes on five datasets are reported in Table 3. The experimental results show that, first, IiAGLMC exhibits the shortest runtime and lowest algorithmic complexity across these datasets. Second, RMSC achieves fast execution and short runtime due to its low computational complexity. Furthermore, LT-MSC, which clusters using the matrix nuclear norm, has a relatively low algorithmic complexity but a relatively long runtime. In comparison, ETSPMSC's runtime is moderate, primarily due to the increased runtime required to process more data and perform more computational steps. Despite this, ETSPMSC's algorithmic complexity is similar to that of t-SVD-MSC, but its runtime speed and efficiency are significantly superior to t-SVD-MSC.

[0165] Table 3 Running time of multi-view clustering algorithm

[0166] Table3 The running time of multi view clustering algorithm

[0167]

[0168]

[0169] 2.6 Convergence Analysis

[0170] The convergence curves of the ETSPMSC algorithm on three benchmark datasets, COIL-20, UCI, and Prokaryotic, are plotted. Figure 17 、 Figure 18 as well as Figure 19 It can be seen that the objective function can quickly approach zero around the 20th iteration, which indicates that the algorithm can enter a stable convergence stage after a finite number of iterations.

[0171] Example 2

[0172] This embodiment discloses a multi-view subspace clustering system for enhancing tensor Schatten P-norm, which can implement the method of the above embodiment, including a denoising model construction module, an ETSPMSC model construction module, and an ETSPMSC model optimization module;

[0173] The denoising model construction module sets multi-view data; constructs a similarity matrix of the view data set to be processed; then constructs a probability transfer matrix based on the similarity matrix, and constructs an initial denoising model in conjunction with the probability transfer matrix;

[0174] The ETSPMSC model construction module applies the weighted tensor Schatten P-norm to the initial denoising model to obtain an initial ETSPMSC model;

[0175] The ETSPMSC model optimization module cooperates with the multi-view data and adopts the alternating direction multiplier method to optimize the initial ETSPMSC model to obtain the optimized ETSPMSC model.

[0176] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0177] The preferred embodiments of the invention disclosed above are intended only to help illustrate the invention. These preferred embodiments do not exhaust all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. Enhanced tensor Schatten P-norm multi-view subspace clustering method, characterized by: The following steps are involved: S1. Constructing an initial ETSPMSC model; specifically comprising the following steps: S11, setting multi-view data; constructing a similarity matrix using the similarity between multi-view data points; and then constructing a probability transfer matrix based on the similarity matrix; S12, constructing an initial denoising model in conjunction with the probability transfer matrix; the initial denoising model includes Laplace noise and Gaussian noise; S13, applying the weighted tensor Schatten P-norm to the initial denoising model to obtain an initial ETSPMSC model; S2, updating and adjusting the variables in the initial ETSPMSC model using the alternating direction multiplier method in conjunction with the multi-view data to optimize the initial ETSPMSC model to obtain an optimized ETSPMSC model; S3. Using the optimized ETSPMSC model to perform denoising on the view dataset to be processed, to obtain a denoised view tensor; S4. Stack the denoised view tensors into a matrix and perform clustering on the matrix using spectral clustering.

2. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 1, characterized in that: The S12 includes the following steps: S121, decomposing the probability transfer matrix into a probability matrix and an error matrix; S122, applying a low-rank constraint to the probability matrix to obtain a probability matrix after the low-rank constraint; S123, stack the probability matrices in S122 along the direction of modulo 3 to generate a probability matrix tensor, and explore various noise pollutions from a high-dimensional level; S124. Considering different noise information in the data at the same time, the error matrix in S121 is deleted, and an initial denoising model is constructed to characterize and remove noise pollution.

3. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 2, characterized in that: The low-rank constraint formula is as follows: Where: E v represents the error matrix, P v represents the probability transfer matrix, Z v represents the probability matrix, ||·|| * represents the low-rank constraint, ||·|| 2,1 Indicates l 2,1 -norm constraint, V represents the total number of views, and λ is a smoothness term.

4. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 3, characterized in that: The initial denoising model formula described in S123 is as follows: Where: E represents the Gaussian noise tensor, N represents the Laplace noise tensor; P is the probability transfer tensor; Z is the probability tensor, which is formed by stacking the probability matrices along the modulo 3 direction; λ1 and λ2 are regularization parameters used to constrain E and N, ||·|| TNN is the tensor core norm constraint, and ||·||1 is the l1-norm constraint of the tensor.

5. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 4, characterized in that: The specific process of S13 applying the weighted tensor Schatten P-norm to the initial denoising model is as follows: The weighted tensor Schatten P-norm is used to replace the TNN in the initial denoising model, and a graph regularization constraint is imposed on the view to ensure the local structural characteristics of the view. Where tr(·) represents the trace of the matrix, L v is the Laplace matrix, represents the weighted tensor Schatten P-norm, and λ3 is the regularization parameter.

6. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 5, characterized in that: The S2 comprises the following steps: S21, setting an auxiliary tensor; using the auxiliary tensor to perform equivalent transformation on the initial ETSPMSC model to obtain a transformed ETSPMSC model; S22, solving the transformed ETSPMSC model using an alternating direction multiplier method to obtain a solved ETSPMSC model; S23. Using an alternating direction multiplier method to update the variables in the solved ETSPMSC model.

7. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 6, characterized in that: The S23 includes the following steps: S231, setting adjustable parameters; inputting the multi-view data and the adjustable parameters into the solved ETSPMSC model; S232, initializing the parameters in the solved ETSPMSC model to obtain an initialized ETSPMSC model; S233, setting a convergence formula; using the alternating direction multiplier method to calculate the solution formula of multiple variables in the initialized ETSPMSC model; and then updating the multiple variables in the initialized ETSPMSC model in conjunction with the solution formula of the multiple variables until the convergence formula is established.

8. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 7, characterized in that: The multiple variables in the initialized ETSPMSC model in S233 include a probability tensor, an auxiliary tensor, a Gaussian noise tensor, a Laplace noise tensor, a Lagrange multiplier, and a penalty parameter.

9. The multi-view subspace clustering method based on the enhanced tensor Schatten P-norm according to claim 8, characterized in that: The convergence formula is as follows, ||P-Z-E-N|| ∞ <ε,||Z-T|| ∞ <ε(4) Among them, ||PZEN|| ∞ <ε and ||ZT|| ∞ <ε indicates that the infinite norm of the residual should be less than the threshold ε; this indicates that the decomposition result is close enough to the original data and the error is within the controllable range; ε=10 -7 , when the convergence condition is less than ε, the algorithm is considered to have converged; T represents the auxiliary tensor; ||·|| ∞ is the infinity norm of the tensor.

10. A system for implementing the multi-view subspace clustering method for enhancing the tensor Schatten P-norm according to any one of claims 1 to 9.