A robust fuzzy subspace image clustering method based on low-dimensional representation
Patent Information
- Application Number
- CN202410312765.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-19
AI Technical Summary
[0006]为了避免现有技术的不足之处,本发明提出一种基于低维表示的鲁棒模糊子空间图像聚类方法,针对现有的均值类聚类算法难以处理高维数据的问题,为了研究数据的低维子空间结构,以期在一个鲁棒低维子空间中对数据完成聚类任务
[0036]本发明提出的一种基于低维表示的鲁棒模糊子空间图像聚类方法,其有益效果具体包括:
Smart Images

Figure CN118154916B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image recognition and classification and pattern recognition, and relates to a robust fuzzy subspace image clustering method based on low-dimensional representation. Background Technology
[0002] Unsupervised clustering algorithms group data samples based on their similarity or dissimilarity, grouping highly similar samples into one cluster and less similar samples into different clusters. Unsupervised clustering can effectively discover natural groupings within data. With the development of information technology and the internet, the ability to collect and generate data has continuously improved, and the types and feature dimensions of data have become increasingly diverse and complex. Against this backdrop, researchers are eager to extract useful information from high-dimensional, complex data. To this end, some researchers have proposed a series of robust subspace clustering algorithms, aiming to reduce redundant information and noise in high-dimensional data while completing the clustering task. These algorithms have been widely applied in important fields such as facial recognition and medical image analysis.
[0003] The advantage of subspace clustering algorithms lies in their ability to leverage the correlations between multiple data dimensions, projecting the dataset into a lower-dimensional subspace and performing clustering within that subspace. This reduces data dimensionality, eliminates redundant information, decreases computational complexity, and improves the accuracy and robustness of clustering. In summary, unsupervised clustering and subspace clustering algorithms have significant application value in data mining, pattern recognition, and machine learning, helping us extract meaningful information from large amounts of complex data and reveal the inherent relationships between data samples.
[0004] Gao Haiyan et al. (Robust Adaptive Symmetric Nonnegative Matrix Factorization Clustering Algorithm [J]. Computer Applications Research, 2023, 40(04): 1024-1029.) proposed a robust adaptive symmetric nonnegative matrix factorization clustering algorithm. This algorithm aims to reduce the dimensionality and cluster the data by performing nonnegative matrix factorization, representing the data as the product of two nonnegative matrices. It utilizes the sensitivity of nonnegative matrix factorization to initialization features to gradually enhance clustering performance, while also leveraging L... 2,1Norms mitigate the impact of noise and outliers, maintain feature rotation invariance, and improve the robustness of the algorithm model. However, while this algorithm improves stability and robustness, the model is relatively redundant, has low practical application value, remains difficult for high-dimensional data processing, and its actual performance is worse than that of this invention. For example, when processing the USPS image dataset with 9298 samples and 256 dimensions, the method proposed by Gao Haiyan et al. could not effectively remove redundant information and establish accurate indexes for such high-dimensional datasets, affecting the accuracy and efficiency of data processing results. Summary of the Invention
[0005] Technical problems to be solved
[0006] To avoid the shortcomings of existing technologies, this invention proposes a robust fuzzy subspace image clustering method based on low-dimensional representation. This method addresses the problem that existing mean-based clustering algorithms struggle to handle high-dimensional data. In order to study the low-dimensional subspace structure of the data, the aim is to complete the clustering task in a robust low-dimensional subspace.
[0007] Technical solution
[0008] A robust fuzzy subspace image clustering method based on low-dimensional representation, characterized by the following steps:
[0009] Step 1: Scale the dataset containing n images into a data matrix. Where d = a × b is the dimension of the data sample, n is the number of data samples, and x i This represents the i-th sample;
[0010] The image is a×b pixels in size;
[0011] Step 2: For a sample data matrix X, the low-dimensional representation of the data is to find two matrices to approximate the data matrix, that is, to represent the data matrix in the form X≈UV;
[0012] Random assignment indicator matrix And make it satisfy U T U = I, initialize the membership matrix and low-dimensional representation matrix V∈R m×n Given the number of normal sample points k, the objective function of the robust blurred subspace image clustering method based on low-dimensional representation is as follows:
[0013]
[0014] Step 3: Calculate the weight vector s based on the number of normal sample points k.
[0015] Step 4: Update and calculate the low-dimensional cluster center matrix M, where m jThe column vectors of the cluster center matrix:
[0016]
[0017] Update the optimal mean vector b of the calculated data:
[0018]
[0019] Step 5: Calculate the indicator matrix U based on the singular value decomposition (SVD):
[0020]
[0021] in, It is an orthogonal matrix, E T E = I, It is a diagonal matrix. It is an orthogonal matrix, F T F = I;
[0022] Based on the values obtained from initialization and update, the indicator matrix U in the low-dimensional representation of the data is calculated using the following expression.
[0023] U = E[I m ,0]F T
[0024] Step 6: Calculate the low-dimensional representation matrix V, v of the data. i These are the row vectors of matrix V:
[0025]
[0026] Step 7: Calculate the membership matrix Y, where y ij It is the j-th element in the i-th row of Y:
[0027]
[0028] Step 8: Repeat steps 3 through 7, iteratively updating the elements in each step until the value of the objective function converges;
[0029] The membership matrix Y after the objective function converges and the sample matrix V after low-dimensional representation are finally obtained through the alternating optimization in step 8. The column containing the maximum value of each row in the membership matrix Y is selected to obtain the final clustering result.
[0030] When calculating the weight vector s, the s corresponding to the first k normal point samples i The remainder is 1, and the remaining s i The value is 0; s i When = 1, the corresponding sample point is a normal point; s iWhen = 0, the corresponding sample point is an outlier; for normal sample points, their corresponding objective function values are calculated normally, while the function value corresponding to outlier points is 0, and the calculation only considers s. i =1 when the sample point is a normal value.
[0031] The solution to s is: s corresponding to the normal point sample. i =1, the discrete value corresponding to s i =0.
[0032] A computer program product characterized by including computer-executable instructions, which, when executed, are used to implement the robust fuzzy subspace image clustering method based on low-dimensional representation.
[0033] An electronic device, characterized in that it includes a processor and a memory, the processor being configured to implement the steps of the robust fuzzy subspace image clustering method based on low-dimensional representation when executing a computer program stored in the memory.
[0034] A readable storage medium, characterized in that a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, it implements the data steps of the robust fuzzy subspace image clustering method based on low-dimensional representation.
[0035] Beneficial effects
[0036] The present invention proposes a robust fuzzy subspace image clustering method based on low-dimensional representation, the specific benefits of which include:
[0037] (1) This invention uses the low-dimensional representation of data to replace the classical data dimensionality reduction method to obtain a low-dimensional subspace, thereby reducing the loss of original information during projection, better preserving the important features of the data, and considering the discriminability of the subspace, it can obtain a discriminative low-dimensional subspace.
[0038] (2) In the low-dimensional representation process, the present invention introduces an optimal mean vector to remove the mean value of the data and remove the offset value of the data, which can enhance the robustness of the method to outliers and noise, and make the clustering results more stable and reliable.
[0039] (3) The present invention introduces new weight vectors. By using these weight vectors, it can be ensured that the results of outliers removed during the low-dimensional representation and clustering process are consistent in a statistical sense, thereby improving the robustness of both low-dimensional representation and clustering. Attached Figure Description
[0040] Figure 1 Flowchart of a robust fuzzy subspace image clustering method based on low-dimensional representation
[0041] Figure 2Flowchart of the specific implementation on the USPS dataset Detailed Implementation
[0042] The present invention will now be further described in conjunction with the embodiments and accompanying drawings:
[0043] A robust fuzzy subspace image clustering method based on low-dimensional representation, characterized by the following steps:
[0044] Step 1: Scale the dataset containing n a×b pixels into a single data matrix. Where d = a × b is the dimension of the data sample, n is the number of data samples, and x i This represents the i-th sample;
[0045] Step 2: Based on the image data matrix X, randomly assign values to the indicator matrix. And make it satisfy U T U = I, initialize the membership matrix and low-dimensional representation matrix V∈R m×n Given the number of normal sample points k;
[0046] Step 3: Calculate the weight vector s based on the initialization and the number of normal sample points k;
[0047] Step 4: Based on the initialization, update the low-dimensional cluster center matrix M using the following expression.
[0048]
[0049] Step 5: Based on the low-dimensional cluster center matrix M obtained from initialization and updating, calculate the optimal mean vector b of the data using the following expression.
[0050]
[0051] Step 6: According to singular value decomposition, we can obtain:
[0052]
[0053] in, It is an orthogonal matrix, E T E = I, It is a diagonal matrix. It is an orthogonal matrix, F T F = I. Based on the values obtained from initialization and update, the basis matrix U in the low-dimensional representation of the data is calculated using the following expression.
[0054] U = EF T (4)
[0055] Step 7: Based on the values obtained from initialization and update, calculate the vector representation of the low-dimensional representation matrix V of the data using the following expression:
[0056]
[0057] Step 8: Based on the values obtained from initialization and update, calculate the membership matrix Y using the following expression:
[0058]
[0059] Step 9: Through the alternating optimization in Step 8, we can obtain the membership matrix Y after the objective function converges and the sample matrix V after low-dimensional representation. Selecting the column containing the maximum value for each row of the membership matrix Y yields the final clustering result. Detailed Implementation
[0061] This invention proposes a robust fuzzy subspace image clustering method based on low-dimensional representation. The specific implementation steps are illustrated using the USPS dataset as an example, but the technical content of this invention is not limited to the scope described herein. The selected USPS dataset contains 9298 samples, with an original data dimension of 256 and a total of 10 classes.
[0062] Implementation Step 1: Stretch the 9298 samples into a data matrix
[0063] Implementation Step 2: Construct a fuzzy subspace image clustering model based on low-dimensional representation as follows:
[0064]
[0065] For a given sample data matrix X, the low-dimensional representation of the data aims to find two matrices to approximate the data matrix, that is, to represent the data matrix in the form X≈UV. It is an indicator matrix. This is a low-dimensional representation matrix. Furthermore, it is desirable to impose an orthogonal constraint U on the indicator matrix. T U = I is used to obtain a unique low-dimensional representation of the data matrix. Based on the obtained low-dimensional representation matrix, the goal is to directly perform fuzzy clustering on the low-dimensional data. r is the fuzzy coefficient, v... i For the i-th sample in the low-dimensional representation, m j For the j-th cluster center, Let y represent the membership matrix. ij Let λ represent the probability of the i-th low-dimensional sample belonging to the j-th cluster, and λ be the regularization parameter. For identity matrix, 1 10 It is a 10-dimensional vector of all ones.
[0066] To improve the clustering performance of the model in noisy environments and enhance its robustness, a robust blurred image clustering model based on low-dimensional representation is constructed as follows:
[0067]
[0068] Among them, 1 9298 It is a 9298-dimensional all-one vector, which increases the amount of data x compared to the initial model. i bias vector This is used to remove the mean, which improves the robustness of the method. A weight vector is also introduced. And give it The constraint is k, which is the number of normal sample points. This model introduces the same weight vector at both ends of the model, which can filter the sample points and assign a weight of 0 to the outliers, so that the outliers do not participate in the calculation of the low-dimensional representation and clustering, thus achieving the effect of removing outliers in both parts at the same time.
[0069] Implementation Step 3: For the weight vector To solve for the objective function (8), it is transformed into the following form:
[0070]
[0071] in, To facilitate subsequent solutions. Because s i The solution is a discrete value, with the first k normal point samples corresponding to a value of 1, and the rest being 0. i When = 1, the corresponding sample point is considered a normal point. i When = 0, the corresponding sample point is considered an outlier. Normal sample points are calculated normally for their corresponding objective function values; the function value corresponding to an outlier is 0, therefore subsequent calculations only consider s. i =1 is the case when the sample point is a normal value.
[0072] Implementation Step 4: For the subspace cluster center matrix The solution can be obtained by directly taking the partial derivative of the objective function (8):
[0073]
[0074] Implementation Step 5: For the optimal mean vector Solving this equation, the original objective function (8) is transformed into:
[0075]
[0076] Furthermore, we can transform it into the form of a matrix trace for easier solution:
[0077]
[0078] Taking the derivative of b directly and setting it to 0, we get:
[0079]
[0080] Implementation Step 6: For the indicator matrix Solving this equation, the original objective function (8) is transformed into:
[0081]
[0082] According to singular value decomposition, we can obtain:
[0083]
[0084] in, It is an orthogonal matrix, E T E = I, It is a diagonal matrix. It is an orthogonal matrix, F T F = I. Based on the results of singular value decomposition, the optimal expression for U is:
[0085] U = E[I m ,0]F T (16)
[0086] Implementation Step 7: For low-dimensional representation vectors Solving for the objective function (8), the objective function is transformed into:
[0087]
[0088] in, It is by The diagonal matrix formed It is by The matrix formed. The above equation can be directly applied to v. i Differentiation yields:
[0089]
[0090] Implementation Step 8: For the membership matrix In solving this problem, the objective function (8) now has only one variable y with a power higher than 1. ij By applying the Lagrange multiplier method, we can easily obtain:
[0091]
[0092] Implementation step 9: Repeat steps 3 to 8 until the objective function (8) converges, and output the membership matrix. By selecting the column containing the maximum value in each row of the membership matrix Y, we obtain the label matrix F, which yields the final image clustering result.
[0093] Finally, comparing the obtained label matrix F with the true labels of the samples, the clustering accuracy (ACC) of this invention on the USPS digital image dataset was 70.73%, and the clustering normalized mutual information (NMI) was 68.24%, which is a significant improvement compared to the comparison algorithm. This also experimentally verifies the effectiveness of this invention in image clustering. Therefore, although the robust fuzzy subspace image clustering method proposed in this invention removes a large number of data features from the original image data, it not only improves the clustering accuracy of face image data but also significantly reduces the data size, verifying the effectiveness of this invention.
[0094] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the scope of the technology disclosed in the present invention, and such modifications or substitutions should all be covered within the scope of protection of the present invention.
Claims
1. A robust fuzzy subspace image clustering method based on low-dimensional representation, characterized in that... The steps are as follows: Step 1: Include The dataset of images is stretched into a data matrix. ,in For the dimensions of the data samples, The number of data samples. Indicates the first One sample; The image is Pixel scale; Step 2: For a sample data matrix The low-dimensional representation of data involves finding two matrices to approximate the data matrix, that is, representing the data matrix as... The form; Random assignment indicator matrix and make it satisfy Initialize the membership matrix and low-dimensional representation matrix And give the number of normal sample points. The objective function for constructing a robust fuzzy subspace image clustering method based on low-dimensional representation is as follows: Step 3: Based on the number of normal sample points The weight vector is calculated. , ; Step 4: Update and calculate the low-dimensional cluster center matrix ,in The column vectors of the cluster center matrix: Update the optimal mean vector of the calculated data. : Step 5: Calculate the indicator matrix based on Singular Value Decomposition (SVD). : in, It is an orthogonal matrix. , It is a diagonal matrix. It is an orthogonal matrix. ; Based on the values obtained from initialization and update, the indicator matrix in the low-dimensional representation of the data is calculated using the following expression. Step 6: Calculate the low-dimensional representation matrix of the data , yes Row vectors of a matrix: Step 7: Calculate the membership matrix ,in yes The Line 1 One element: Step 8: Repeat steps 3 through 7, iteratively updating the elements in each step until the value of the objective function converges; The membership matrix after the objective function converges is finally obtained through the alternating optimization in step 8. and the sample matrix after low-dimensional representation ; for membership matrix Select the column containing the maximum value for each row to obtain the final clustering result.
2. The robust fuzzy subspace image clustering method based on low-dimensional representation according to claim 1, characterized in that: The calculated weight vector At that time, before The corresponding normal point samples =1, the remainder The value is 0; At that time, the corresponding sample point is a normal point; When this occurs, the corresponding sample point is an outlier; For normal sample points, their corresponding objective function values are calculated normally. For outlier points, the function value is 0, and the calculation only considers... The case where the sample points are normal values.
3. The robust fuzzy subspace image clustering method based on low-dimensional representation according to claim 2, characterized in that: The The solution is: the normal point sample corresponding to The discrete value corresponding to .
4. A computer program product, characterized in that... It includes computer-executable instructions, which, when executed, are used to implement the method described in any one of claims 1 to 3.
5. An electronic device, characterized in that, It includes a processor and a memory, wherein the processor is configured to implement the steps of the robust fuzzy subspace image clustering method based on low-dimensional representation as described in any one of claims 1 to 3 when executing a computer program stored in the memory.
6. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the robust fuzzy subspace image clustering method based on low-dimensional representation as described in any one of claims 1 to 3.