A multi-view subspace clustering method, device, equipment and storage medium

CN114612671BActive Publication Date: 2025-05-23HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210158539.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-05-23
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

The existing multi-view subspace clustering method lacks robustness when processing multi-view data, making it difficult to explore high-order cross-view correlation, and has a large space-time overhead, making it difficult to apply to massive multi-view clustering tasks.

Method used

A one-step tensor low-rank method is proposed. By performing feature extraction, self-representation processing, tensor singular value decomposition and affinity matrix calculation on multi-view data, subspace clustering is used to improve the robustness and accuracy of the method.

Benefits of technology

Through the one-step tensor low-rank method, the robustness and accuracy of the multi-view subspace clustering method can be effectively improved, and the influence of noise and outliers can be reduced, which is suitable for massive multi-view clustering tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612671B_ABST
    Figure CN114612671B_ABST
Patent Text Reader

Abstract

The present embodiment provides a multi-view subspace clustering method, device, equipment and storage medium, which belongs to the field of pattern recognition technology. The method includes: extracting features from multi-view data to obtain a data feature matrix; performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view data; constructing a first representation tensor of the multi-view data according to the self-representation matrix; performing tensor singular value decomposition on the first representation tensor to obtain a second representation tensor; calculating the affinity matrix of the multi-view data based on the second representation tensor; and using a spectral clustering algorithm to segment the affinity matrix to obtain a subspace clustering result. The present application can cluster multi-view data through a one-step tensor low-rank method to improve the robustness and accuracy of the multi-view subspace clustering method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of pattern recognition technology, and in particular to a multi-view subspace clustering method, device, equipment and storage medium. Background Art

[0002] Subspace clustering aims to divide samples into different clusters and find a low-dimensional subspace representation based on their similarities without label information. It consists of two steps: first, construct an affinity matrix to describe the relationship between multimedia data points; then, apply a clustering algorithm on the affinity matrix to obtain the final clustering result. Therefore, the quality of the affinity matrix determines the clustering performance to a large extent, but due to the presence of noise and outliers, the affinity matrix constructed on the original data features is often not robust enough and the correlation of multi-view data is often not fully mined.

[0003] Existing multi-view subspace clustering methods mainly carry out two-step learning, namely learning the representation tensor and learning the affinity matrix. In the two independent steps, the affinity matrix is ​​fixedly solved according to the representation tensor, which cannot effectively mine the high correlation between the two; at the same time, the tensor representation learning method of multi-view data lacks robustness, is easily affected by noise and outliers, and is difficult to explore high-order cross-view correlations. In the graph-based multi-view subspace clustering algorithm, the graph is often directly constructed according to the learned self-representation matrix, which lacks flexibility, resulting in large time and space overhead, and makes the existing methods difficult to apply to massive multi-view clustering tasks. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to propose a multi-view subspace clustering method, apparatus, device and storage medium, which can cluster multi-view data through a one-step tensor low-rank method to improve the robustness and accuracy of the multi-view subspace clustering method.

[0005] To achieve the above object, a first aspect of an embodiment of the present application proposes a multi-view subspace clustering method, comprising:

[0006] Extract features from multi-view data to obtain a data feature matrix;

[0007] Performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view data;

[0008] constructing a first representation tensor of the multi-view data according to the self-representation matrix;

[0009] Performing tensor singular value decomposition on the first representation tensor to obtain a second representation tensor;

[0010] Obtaining an affinity matrix of the multi-view data by calculation based on the second representation tensor;

[0011] The affinity matrix is ​​segmented using a spectral clustering algorithm to obtain a subspace clustering result.

[0012] In some embodiments, performing self-representation processing on the data feature matrix to obtain the self-representation matrix of the multi-view data includes:

[0013] Get the data feature matrix X of the multi-view data v ;

[0014] The data feature matrix is ​​processed by self-representation using the following formula to obtain the self-representation matrix Z of the multi-view data: v :

[0015] X v =X v Z v +E v , v=1,2,…,V;

[0016] Wherein, V represents the number of views in the multi-view data, X v represents the data feature matrix in the vth view, Z v represents the self-representation matrix in the v-th view, E v represents the noise matrix in the v-th view.

[0017] In some embodiments, constructing a first representation tensor of the multi-view data according to the self-representation matrix comprises:

[0018] The obtained V self-representation matrices (Z 1 ,Z 2 ,…,Z V ) were used as frontal slices;

[0019] The first representation tensor of the multi-view data is constructed using the following formula:

[0020]

[0021] Wherein, Φ(·) is used to merge the V self-representation matrices as the frontal slices to construct the first representation tensor of the multi-view data

[0022] In some embodiments, performing a tensor singular value decomposition on the first representation tensor to obtain a second representation tensor includes:

[0023] Get the first representation tensor

[0024] The first tensor is represented by the following formula Perform low-rank constraint processing to obtain a second representation tensor of the multi-view data

[0025]

[0026]

[0027] Wherein, the first represents a tensor After tensor singular value decomposition, it is defined as the product of three matrix tensors, represents the first orthogonal tensor, v represents the second orthogonal tensor, v T represents the transpose of the second orthogonal tensor, and represents the diagonal tensor consisting of tensor eigenvalues, n 1 、n 2 、n 3 Respectively represent the first representation tensor The three dimension values ​​of .

[0028] In some embodiments, the calculating the affinity matrix of the multi-view data based on the second representation tensor includes:

[0029] The optimization objective function of the multi-view data is calculated using the following formula:

[0030]

[0031] xT v =X v Z v +E v , v=1,2,…,V;

[0032]

[0033] E=[E 1 ; E 2 ;…;E V ], A T 1=1,0≤A≤1;

[0034] in, represents the first representation tensor, E represents the noise matrix composed of V views, A represents the affinity matrix, represents the nuclear norm, V represents the number of input views, α represents the first penalty parameter, tr(·) represents the trace of the solution matrix, and L A represents the graph Laplacian matrix of the affinity matrix, st represents the constraints that the optimization objective function needs to satisfy, and Zv represents the self-representation matrix in the vth view, β represents the second penalty parameter corresponding to the affinity matrix, γ represents the third penalty parameter corresponding to the noise matrix, X v represents the data feature matrix in the v-th view;

[0035] The affinity matrix of the multi-view data is obtained according to the optimization objective function.

[0036] In some embodiments, obtaining the affinity matrix of the multi-view data according to the optimization objective function includes:

[0037] Using the alternating direction multiplier method to solve the optimization parameters of the optimization objective function;

[0038] An affinity matrix of the multi-view data is solved according to the optimization parameters.

[0039] A second aspect of the embodiments of the present application provides a multi-view subspace clustering device, comprising:

[0040] A feature extraction module is used to extract features from multi-view data to obtain a data feature matrix;

[0041] A self-representation processing module, used for performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view data;

[0042] A representation tensor construction module, configured to construct a first representation tensor of the multi-view data according to the self-representation matrix;

[0043] A tensor singular value decomposition module, used for performing tensor singular value decomposition on the first representation tensor to obtain a second representation tensor;

[0044] an affinity matrix calculation module, configured to calculate an affinity matrix of the multi-view data based on the second representation tensor;

[0045] The subspace clustering module is used to segment the affinity matrix using a spectral clustering algorithm to obtain a subspace clustering result.

[0046] In some embodiments, the tensor singular value decomposition module is used to perform tensor singular value decomposition on the first representation tensor to obtain a second representation tensor, including:

[0047] Get the data feature matrix X of the multi-view data v ;

[0048] The data feature matrix is ​​processed by self-representation using the following formula to obtain the self-representation matrix Z of the multi-view data: v :

[0049] X v =X v Z v +E v , v=1,2,…,V;

[0050] Wherein, V represents the number of views in the multi-view data, X v represents the data feature matrix in the vth view, Z v represents the self-representation matrix in the v-th view, E v represents the noise matrix in the v-th view.

[0051] A third aspect of an embodiment of the present application proposes a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor is used to execute a multi-view subspace clustering method as described in any one of the embodiments of the first aspect of the present application.

[0052] The fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a computer, the computer is used to execute a multi-view subspace clustering method as described in any one of the embodiments of the first aspect of the present application.

[0053] The embodiment of the present application proposes a multi-view subspace clustering method, device, equipment and storage medium, which extracts features from multi-view data to obtain a data feature matrix; performs self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view data; constructs a first representation tensor of the multi-view data according to the self-representation matrix; performs tensor singular value decomposition on the first representation tensor to obtain a second representation tensor; calculates an affinity matrix of the multi-view data based on the second representation tensor; and uses a spectral clustering algorithm to segment the affinity matrix to obtain a subspace clustering result. The present application can cluster multi-view data through a one-step tensor low-rank method to improve the robustness and accuracy of the multi-view subspace clustering method. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a flowchart of a multi-view subspace clustering method provided in an embodiment of the present application;

[0055] Figure 2 is a schematic diagram of a tensor singular value decomposition of a first representation tensor provided in an embodiment of the present application;

[0056] Figure 3 is a flowchart of step S152 in an embodiment of the present application;

[0057] Figure 4It is a process architecture diagram of a multi-view subspace clustering method provided by a specific embodiment of the present application;

[0058] Figure 5 It is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0060] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0062] First, some nouns involved in this application are analyzed:

[0063] Artificial intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. AI is a branch of computer science. AI attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. AI can simulate the information process of human consciousness and thinking. AI is also a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0064] Spectral Clustering (SC): It is based on spectral graph theory. Compared with traditional clustering algorithms, it has the advantages of being able to cluster in sample spaces of arbitrary shapes and converge to the global optimal solution. The main idea of ​​this algorithm is to regard each object in the data set as a vertex of a graph. These vertices can be connected by edges, and the similarity between the vertices is quantified as the weight of the edge connecting the corresponding vertices. The weight of the edge between two points that are farther away is lower, and the weight of the edge between two points that are closer is higher. Then, by cutting the graph composed of all data points, the sum of the weights of the edges between different subgraphs after cutting is as low as possible, and the sum of the weights of the edges within the subgraph is as high as possible, so as to achieve the purpose of clustering.

[0065] Tensor Singular Value Decomposition (t-SVD): It is generated based on tube fiber convolution. It can not only fully express the correlation in spatial structure more than other tensor decomposition methods, but also can perform fast calculations through Fourier transform to improve computational efficiency.

[0066] The Alternating Direction Method of Multipliers (ADMM) is a computational framework for solving separable convex optimization problems. Due to its fast processing speed and good convergence performance, ADMM is suitable for solving distributed convex optimization problems, especially statistical learning problems.

[0067] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0068] The multi-view subspace clustering method provided in the embodiment of the present application can be applied to artificial intelligence. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0069] Subspace clustering aims to divide samples into different clusters and find a low-dimensional subspace representation based on their similarities without label information. It consists of two steps: first, construct an affinity matrix to describe the relationship between multimedia data points; then, apply a clustering algorithm on the affinity matrix to obtain the final clustering result. Therefore, the quality of the affinity matrix determines the clustering performance to a large extent, but due to the presence of noise and outliers, the affinity matrix constructed on the original data features is often not robust enough and the correlation of multi-view data is often not fully mined.

[0070] Existing multi-view subspace clustering methods mainly carry out two-step learning, namely learning the representation tensor and learning the affinity matrix. In the two independent steps, the affinity matrix is ​​fixedly solved according to the representation tensor, which cannot effectively mine the high correlation between the two; at the same time, the tensor representation learning method of multi-view data lacks robustness, is easily affected by noise and outliers, and is difficult to explore high-order cross-view correlations. In the graph-based multi-view subspace clustering algorithm, the graph is often directly constructed according to the learned self-representation matrix, which lacks flexibility, resulting in large time and space overhead, and makes the existing methods difficult to apply to massive multi-view clustering tasks.

[0071] Based on this, the main purpose of the embodiments of the present application is to propose a multi-view subspace clustering method, device, equipment and storage medium, which can cluster multi-view data through a one-step tensor low-rank method to improve the robustness and accuracy of the multi-view subspace clustering method.

[0072] A multi-view subspace clustering method provided in an embodiment of the present application can be applied to a terminal, can be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc.; the server can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the above method, etc., but is not limited to the above forms.

[0073] Embodiments of the present application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0074] Reference Figure 1 According to a multi-view subspace clustering method of the first aspect of the embodiments of the present application, it includes but is not limited to steps S110 to S160.

[0075] S110, extracting features from the multi-view data to obtain a data feature matrix;

[0076] S120, performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view data;

[0077] S130, constructing a first representation tensor of multi-view data according to the self-representation matrix;

[0078] S140, performing tensor singular value decomposition on the first representation tensor to obtain a second representation tensor;

[0079] S150, calculating an affinity matrix of the multi-view data based on the second representation tensor;

[0080] S160, using a spectral clustering algorithm to segment the affinity matrix to obtain a subspace clustering result.

[0081] In step S110, feature extraction is performed on the multi-view data to obtain a data feature matrix. Specifically, the multi-view data is composed of different representations of the same data. In order to better describe the multi-modal features of the view image, a multi-view matrix composed of V views is input. Among them, X v represents the data feature matrix of the vth view among V views, d v represents the dimension of the vth data feature matrix, and n represents the number of data points in the vth data feature matrix.

[0082] It should be noted that, in order to facilitate subsequent data processing, each data point in the data feature matrix may be normalized.

[0083] In some embodiments, step S120 specifically includes steps S110 to S120.

[0084] S121, obtaining the data feature matrix X of the multi-view data v ;

[0085] S122, using the following formula (1) to perform self-representation processing on the data feature matrix to obtain the self-representation matrix Z of the multi-view data v :

[0086] X v =X v Z v +E v ,v=1,2,…,V (1)

[0087] Where V represents the number of views in the multi-view data, X v represents the data feature matrix in the vth view, Z v represents the self-representation matrix in the v-th view, E v represents the noise matrix in the v-th view.

[0088] In step S121 to step S122, the data feature matrix is ​​processed for self-representation to obtain the self-representation matrix of the multi-view data. Specifically, since the self-representation of the data depends on the assumption that each data sample can be represented as a linear combination of other data samples in the same subspace, the data feature matrix X of the multi-view data is obtained by v , and the data feature matrix X v Perform self-representation processing, that is, use formula (1) to transform the data feature matrix X v Perform self-representation processing to obtain the self-representation matrix Z of multi-view data v , where the self-representation matrix Z v The (i,j)th element in represents the sample x in the vth view. j In the reconstruction sample x i The contribution made in the process reflects the sample x i and x j The similarity relationship between them.

[0089] In some embodiments, step S130 specifically includes steps S131 to S132.

[0090] S131, the obtained V self-representation matrices (Z 1 ,Z 2 ,…,Z V) were used as frontal slices;

[0091] S132, constructing a first representation tensor of multi-view data using the following formula (2):

[0092]

[0093] Among them, Φ(·) is used to merge V self-representation matrices as frontal slices to construct the first representation tensor of multi-view data

[0094] In step S131 to step S132, in order to mine more underlying structural information of the data and retain the integrity information of the data, tensor representation is used to explore the relationship between different views in the multi-view data, that is, the first representation tensor of the multi-view data is constructed according to the self-representation matrix. Specifically, the obtained V self-representation matrices (Z 1 ,Z 2 ,…,Z V ) are used as front slices respectively, and V front slices are merged according to Φ(·) in formula (2) to construct a third-order representation tensor, that is, the first representation tensor of multi-view data is obtained where n×n represents the matrix size of the view-specific representation tensor and V represents the number of views in the multi-view data.

[0095] In some embodiments, step S140 specifically includes steps S141 to S142.

[0096] S141, obtain the first representation tensor

[0097] S142, using the following formula (3) and formula (4) to represent the first tensor Perform low-rank constraint processing to obtain the second representation tensor of multi-view data

[0098]

[0099]

[0100] Among them, the first represents the tensor After tensor singular value decomposition, it is defined as the product of three matrix tensors, represents the first orthogonal tensor, v represents the second orthogonal tensor, v T represents the transpose of the second orthogonal tensor, and represents the diagonal tensor consisting of tensor eigenvalues, n 1 、n 2 、n 3Respectively represent the first representation tensor The three dimension values ​​of .

[0101] In step S141 to step S142, in order to reduce the influence of noise and outliers in the view, the first representation tensor Perform tensor singular value decomposition, that is, according to formula (3) and formula (4), the first representation tensor The tensor singular value decomposition is used to perform low-rank constraint processing to obtain the second representation tensor Specifically, Figure 2 As shown, get the first representation tensor According to formula (3), after the tensor singular value decomposition, the first representation tensor Defined u, The product form of three matrix tensors, , and v. According to formula (4), by expressing the first tensor Calculate the tensor nuclear norm and get the second representation tensor This can better capture the consistency between multi-view data and view-specific information. In formula (4), Indicates n 3 Each of the frontal slices is n×n (n=min(n 1 ,n 2 )) size, the diagonal elements in the diagonal tensor are the values ​​of the first representation tensor The eigenvalues ​​of the matrix corresponding to the kth frontal slice in Indicates n 3 The sum of the (,i)th element (i.e., the i-th value of the diagonal) of the frontal slice matrix is ​​then performed on n 3 Sum each element of the diagonal matrix corresponding to the diagonal tensor in the frontal slices.

[0102] In some embodiments, step S150 specifically includes steps S151 to S152.

[0103] S151, using the following formula (5) to calculate the optimization objective function of the multi-view data:

[0104]

[0105] in, represents the first representation tensor, E represents the noise matrix composed of V views, A represents the affinity matrix, represents the nuclear norm, V represents the number of input views, α represents the first penalty parameter, tr(·) represents the trace of the solution matrix, and L A represents the graph Laplacian matrix of the affinity matrix, st represents the constraints that need to be satisfied by the optimization objective function, and Zv represents the self-representation matrix in the vth view, β represents the second penalty parameter corresponding to the affinity matrix, γ represents the third penalty parameter corresponding to the noise matrix, X v Represents the data feature matrix in the vth view;

[0106] S152, obtaining an affinity matrix of the multi-view data according to the optimization objective function.

[0107] In step S151 to step S152, in order to learn a more robust affinity matrix, the first penalty parameter α, the second penalty parameter β and the third penalty parameter γ are used to cover various losses and errors, that is, the optimization objective function of the multi-view data is obtained by minimizing formula (5), and then the optimal affinity matrix of the multi-view data can be obtained according to the optimization objective function. In the above optimization objective function, a ij represents the (,j)th item of affinity matrix A. The graph Laplacian matrix of affinity matrix A is represented by L A =D-(A+A T ) / 2, where D is the i-th diagonal item. It should be noted that st(subject to) is used to represent the constraints that need to be satisfied by the optimization objective function. In addition, since the dimensions of different views may be different, the view-specific noise matrices are vertically connected to construct the noise matrix E.

[0108] In some embodiments, reference Figure 3 Step S152 specifically includes but is not limited to steps S210 to S220.

[0109] Step S210, using the alternating direction multiplier method to solve the optimization parameters of the optimization objective function;

[0110] Step S220: solving the affinity matrix of the multi-view data according to the optimization parameters.

[0111] In step S210 to step S220, after solving the optimization parameters of the optimization objective function by using the alternating direction multiplier method, that is, the ADMM optimization algorithm, the affinity matrix A of the multi-view data is solved according to the optimization parameters. Specifically, since formula (4) is obtained by converting the first representation tensor With the objective function and two constraints (X v =X v Z v +E v , v = 1, 2, ..., V and ) coupling correlation to obtain the optimization objective function, and then To achieve decoupling of the optimization objective function, an auxiliary variable is introduced to facilitate the solution. Separate the first representation tensor in the optimization objective function And by fixing other variables and iteratively updating each variable, the The solution process is then solved by solving the auxiliary variables The optimal parameters of the optimization objective function are obtained, and then the affinity matrix of the multi-view data is solved according to the optimized parameters.

[0112] Specifically, the specific process of solving the optimization parameters of the optimization objective function through the ADMM optimization algorithm is as follows.

[0113] Step S310: construct an augmented Lagrangian function based on formula (6): As shown in formula (7);

[0114]

[0115]

[0116] Among them, Θ and Π represent Lagrangian operators, <·,·> represents the inner product, and ρ is the fourth penalty parameter.

[0117] Step S320, fix the other variables in formula (7) to update the first representation tensor Solve for the first representation tensor The optimization subproblem is The optimization problem of the t+1th iteration can be transformed into the formula (8), where t represents the number of iterations.

[0118]

[0119] Step S330, fix the other variables in formula (7) to update the auxiliary variables Solving for auxiliary variables , then for the auxiliary variables The optimization problem of the t+1th iteration can be transformed into the formula (9).

[0120]

[0121] Since different views do not interfere with each other, the V variables Y in formula (9) are v are independent of each other, so formula (9) can be converted into V optimization sub-problems, for example, the variable update problem of the vth view is shown in formula (10).

[0122]

[0123] Therefore, by adding Y v The derivative of is 0, and the solution of formula (10) is shown in formula (11).

[0124]

[0125] Step S340, fix the other variables in formula (7) to update the noise matrix E, solve the optimization sub-problem about the noise matrix E, and introduce a column-wise connection matrix { v} temporary variable Then the optimization problem for the t+1th iteration of the noise matrix E can be transformed into the formula (12).

[0126]

[0127] Therefore, for E in formula (12), t+1 The optimal solution of is shown in formula (13).

[0128]

[0129] Among them, E t+1 (:,j) represents E t+1 The jth column of .

[0130] In step S350, the other variables in formula (7) are fixed to update the affinity matrix A, and the optimization sub-problem about the affinity matrix A is solved. Then, the optimization problem of the affinity matrix A for the t+1th iteration can be transformed into the one shown in formula (14).

[0131]

[0132] Afterwards, using L A =D-(A+A T ) / 2 replaces the graph Laplacian matrix Then formula (14) can be equivalently replaced as shown in formula (15).

[0133]

[0134] sA T 1=1,0≤A≤1 (15)

[0135] Therefore, according to formula (16), the a in the i independent optimization sub-problems in formula (15) is i To solve.

[0136]

[0137] Among them, the variable satisfy Then formula (16) can be converted into the Laplace formula form as shown in formula (17).

[0138]

[0139] Among them, η and δ represent the Laplace multiplication operator. According to the Carlo-Kuhn-Tucker condition, At the same time, in order to ensure that the similarity of sample points within a class is higher than the similarity of sample points between classes, as shown in formulas (18) and (19), the adaptive nearest neighbor method is used to solve the affinity matrix A, that is, only the largest top K items in the affinity matrix A are retained to improve the clustering performance.

[0140]

[0141]

[0142] According to formulas (18) and (19), β in formula (7) is determined by the number K of the adaptive nearest neighbor method.

[0143] In step S360, the other variables in formula (7) are fixed to update the Lagrangian operators v, Π and the fourth penalty parameter ρ, and the optimization subproblem about the Lagrangian operators v, Π and the fourth penalty parameter ρ is solved. Then, the optimization problem of the t+1th iteration of the Lagrangian operators Θ, Π and the fourth penalty parameter ρ can be transformed into the one shown in formula (20).

[0144]

[0145] Among them, the parameters λ>1, ρ max Represents the maximum value of ρ. Therefore, the above formula (8) to formula (20) are used to optimize formula (7), and then the optimal solution of the affinity matrix A is obtained. Afterwards, the obtained affinity matrix A is applied to the spectral clustering algorithm to obtain the subspace clustering result. Therefore, in order to fully explore the multimodality of the data, the present application performs clustering by learning a robust one-step low-rank tensor graph method, that is, by jointly learning the representation tensor and the affinity matrix to solve the problem that the traditional affinity matrix is ​​fixedly solved and cannot explore the high correlation between the two. Specifically, the tensor singular value decomposition is used to perform high-order constraints on the representation tensor to reduce the impact of data noise and outliers, and the K-adaptive nearest neighbor method is used to reconstruct the affinity matrix, thereby improving the efficiency of subspace clustering.

[0146] In a specific embodiment, Figure 4 As shown, feature extraction is performed on multi-view data to obtain a data feature matrix (X 1 ,X 2 ,…,XV ), perform self-representation processing on each data feature matrix to obtain the self-representation matrix of multi-view data (Z 1 ,Z 2 ,…,Z V ), construct the first representation tensor of the multi-view data according to the obtained multiple self-representation matrices In order to better improve the clustering performance of multi-view data and improve the robustness of the representation tensor, the first representation tensor The low-rank constraint processing of the tensor singular value decomposition is performed to obtain the second representation tensor, which is then jointly optimized with the affinity matrix A using the alternating direction multiplier method to solve the optimal affinity matrix A for multi-view data. Finally, the affinity matrix is ​​segmented using a clustering spectral algorithm, and the subspace clustering result is output. This application performs clustering by learning a robust one-step low-rank tensor graph method, that is, by jointly learning the representation tensor and the affinity matrix, to solve the problem that the traditional affinity matrix is ​​fixedly solved and cannot explore the high correlation between the two.

[0147] The embodiment of the present application also proposes a multi-view subspace clustering device, which includes a feature extraction module, a self-representation processing module, a representation tensor construction module, a tensor singular value decomposition module, an affinity matrix calculation module and a subspace clustering module. The feature extraction module is used to extract features from multi-view data to obtain a data feature matrix; the self-representation processing module is used to perform self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view data; the representation tensor construction module is used to construct a first representation tensor of the multi-view data according to the self-representation matrix; the tensor singular value decomposition module is used to perform tensor singular value decomposition on the first representation tensor to obtain a second representation tensor; the affinity matrix calculation module is used to calculate the affinity matrix of the multi-view data based on the second representation tensor; the subspace clustering module is used to segment the affinity matrix using a spectral clustering algorithm to obtain a subspace clustering result. A multi-view subspace clustering device in the embodiment of the present application is used to execute a multi-view subspace clustering method in the above embodiment, and its specific processing process is the same as a multi-view subspace clustering method in the above embodiment, which will not be repeated here.

[0148] In some embodiments, the tensor singular value decomposition module is used to perform a tensor singular value decomposition on the first representation tensor to obtain a second representation tensor, including obtaining a data feature matrix X of the multi-view data. v , and use formula (1) to perform self-representation processing on the data feature matrix to obtain the self-representation matrix Z of the multi-view data v :

[0149] Where V represents the number of views in the multi-view data, X v represents the data feature matrix in the vth view, Zv represents the self-representation matrix in the v-th view, E v represents the noise matrix in the v-th view.

[0150] It should be noted that a multi-view subspace clustering device in the above embodiment of the present application is used to execute a multi-view subspace clustering method in the above embodiment, and its specific processing process is the same as that of a multi-view subspace clustering method in the above embodiment, which will not be repeated here.

[0151] An embodiment of the present application also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor is used to execute a multi-view subspace clustering method as described in any one of the embodiments of the first aspect of the present application.

[0152] Combine the following Figure 5 The hardware structure of the computer device is described in detail. The computer device includes: a processor 501 , a memory 502 , an input / output interface 503 , a communication interface 504 and a bus 505 .

[0153] The processor 501 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0154] The memory 502 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device or RAM (Random Access Memory). The memory 502 can store operating systems and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 502, and the processor 501 is called to execute a multi-view subspace clustering method of the embodiment of the present application;

[0155] Input / output interface 503, used to implement information input and output;

[0156] The communication interface 504 is used to realize the communication interaction between the device and other devices, and the communication can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and the bus 505 is used to transmit information between the various components of the device (such as the processor 501, the memory 502, the input / output interface 503 and the communication interface 504);

[0157] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via a bus 505 .

[0158] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a computer, the computer is used to execute a multi-view subspace clustering method as described in any one of the embodiments of the first aspect of the present application.

[0159] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0160] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0161] It can be understood by those skilled in the art that Figures 1 to 3 The technical solutions shown in the figure do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or a combination of certain steps, or different steps.

[0162] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0164] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0165] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0166] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0167] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0168] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0170] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A multi-view subspace clustering method, It is characterized in that include: Extract features from multi-view image data to obtain a data feature matrix; Performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view image data; constructing a first representation tensor of the multi-view image data according to the self-representation matrix; Performing a tensor singular value decomposition on the first representation tensor to obtain a second representation tensor; wherein the performing a tensor singular value decomposition on the first representation tensor to obtain the second representation tensor includes: obtaining the first representation tensor ; Use the following formula to represent the first tensor Perform low-rank constraint processing to obtain a second representation tensor of the multi-view image data : ; ; Wherein, the first represents a tensor After tensor singular value decomposition, it is defined as the product of three tensors, represents the first orthogonal tensor, represents the second orthogonal tensor, represents the transpose of the second orthogonal tensor, and represents a diagonal tensor consisting of the eigenvalues ​​of the tensor, , , Respectively represent the first representation tensor The three dimension values ​​of ; The affinity matrix of the multi-view image data is calculated based on the second representation tensor; wherein the affinity matrix of the multi-view image data is calculated based on the second representation tensor, comprising: calculating the optimization objective function of the multi-view image data using the following formula: ; V ; ; ; in, represents the first representation tensor, Indicated by The noise matrix composed of views, represents the affinity matrix, represents the tensor nuclear norm, represents the number of views of the input, represents the first penalty parameter, represents the trace of the solution matrix, The graph Laplacian matrix representing the affinity matrix, represents the constraints that the optimization objective function needs to satisfy, Indicates The self-representation matrix in the view, represents the second penalty parameter corresponding to the affinity matrix, represents the third penalty parameter corresponding to the noise matrix, Indicates the The data feature matrix in the views, Indicates the noise matrix in each view; obtaining the affinity matrix of the multi-view image data according to the optimization objective function; The affinity matrix is ​​segmented using a spectral clustering algorithm to obtain a subspace clustering result.

2. A multi-view subspace clustering method according to claim 1, It is characterized in that The step of performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view image data includes: Obtaining a data feature matrix of the multi-view image data ; The data feature matrix is ​​processed by self-representation using the following formula to obtain the self-representation matrix of the multi-view image data: : V ; Among them, V represents the number of views in the multi-view image data, represents the data feature matrix in the th view, represents the self-representation matrix in the th view, represents the noise matrix in the th view.

3. A multi-view subspace clustering method according to claim 2, It is characterized in that The constructing a first representation tensor of the multi-view image data according to the self-representation matrix comprises: Will get Self-representing matrix As frontal slices respectively; The first representation tensor of the multi-view image data is constructed using the following formula: : ; in, For the The self-representation matrices are combined as the front slices to construct the first representation tensor of the multi-view image data. .

4. A multi-view subspace clustering method according to claim 1, It is characterized in that The obtaining the affinity matrix of the multi-view image data according to the optimization objective function comprises: Using the alternating direction multiplier method to solve the optimization parameters of the optimization objective function; An affinity matrix of the multi-view image data is solved according to the optimization parameters.

5. A multi-view subspace clustering device, It is characterized in that include: A feature extraction module is used to extract features from multi-view image data to obtain a data feature matrix; A self-representation processing module, used for performing self-representation processing on the data feature matrix to obtain a self-representation matrix of the multi-view image data; A representation tensor construction module, configured to construct a first representation tensor of the multi-view image data according to the self-representation matrix; A tensor singular value decomposition module is used to perform a tensor singular value decomposition on the first representation tensor to obtain a second representation tensor; wherein the tensor singular value decomposition on the first representation tensor to obtain the second representation tensor includes: obtaining the first representation tensor ; Use the following formula to represent the first tensor Perform low-rank constraint processing to obtain a second representation tensor of the multi-view image data : ; ; Wherein, the first represents a tensor After tensor singular value decomposition, it is defined as the product of three tensors, represents the first orthogonal tensor, represents the second orthogonal tensor, represents the transpose of the second orthogonal tensor, and represents a diagonal tensor consisting of the eigenvalues ​​of the tensor, , , Respectively represent the first representation tensor The three dimension values ​​of ; an affinity matrix calculation module, configured to calculate the affinity matrix of the multi-view image data based on the second representation tensor; wherein the calculating the affinity matrix of the multi-view image data based on the second representation tensor comprises: calculating the optimization objective function of the multi-view image data using the following formula: ; V ; ; ; in, represents the first representation tensor, Indicated by The noise matrix composed of views, represents the affinity matrix, represents the tensor nuclear norm, represents the number of views of the input, represents the first penalty parameter, represents the trace of the solution matrix, The graph Laplacian matrix representing the affinity matrix, represents the constraints that the optimization objective function needs to satisfy, Indicates The self-representation matrix in the view, represents the second penalty parameter corresponding to the affinity matrix, represents the third penalty parameter corresponding to the noise matrix, Indicates the The data feature matrix in the views, Indicates the noise matrix in each view; obtaining the affinity matrix of the multi-view image data according to the optimization objective function; The subspace clustering module is used to segment the affinity matrix using a spectral clustering algorithm to obtain a subspace clustering result.

6. A multi-view subspace clustering device according to claim 5, It is characterized in that The self-representation processing module is used to perform self-representation processing on the data feature matrix to obtain the self-representation matrix of the multi-view image data, including: Obtaining a data feature matrix of the multi-view image data ; The data feature matrix is ​​processed by self-representation using the following formula to obtain the self-representation matrix of the multi-view image data: : V ; in, V represents the number of views in the multi-view image data, Indicates The data feature matrix in the views, Indicates the The self-representation matrix in the view, Indicates the The noise matrix in each view.

7. A computer device, It is characterized in that The computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is used to execute: A multi-view subspace clustering method as claimed in any one of claims 1 to 4.

8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program. When the computer program is executed by a computer, the computer is configured to perform: A multi-view subspace clustering method as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-view-based subspace clustering method, device and equipment and storage medium

    CN109685155A