An Incomplete Multi-View Scene Image Clustering Method Based on Perspective Mutually Exclusive Pseudo-Label Propagation
By adopting a pseudo-label propagation method with mutually exclusive viewpoints and graph convolutional neural networks in multi-view image clustering, the problem of non-complete multi-view image clustering is solved, and a more efficient and accurate clustering effect is achieved.
Patent Information
- Application Number
- CN202311805148.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-12-26
AI Technical Summary
The existing multi-view image clustering method cannot effectively process image data from non-complete view angles, resulting in poor clustering effect and increasing the difficulty of multi-view scene image clustering.
A pseudo-label propagation method based on perspective mutual exclusion is adopted. By extracting pseudo-labels of complete samples and constructing a perspective-specific semi-supervised classifier based on graph convolutional neural networks, pseudo-label propagation and clustering across perspectives is performed.
It improves the clustering accuracy of non-complete multi-view scene images, can effectively process large-scale multi-view data sets, and significantly improves clustering quality and efficiency.
Smart Images

Figure CN117765291B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of scene image clustering processing in image information processing, and particularly relates to an incomplete multi-view scene image clustering method based on perspective exclusive pseudo-label propagation. Background Art
[0002] In the field of image information processing, the development of technologies such as Internet media has brought about a surge in image information and provided convenience for promoting intelligent technology research. Therefore, how to effectively and accurately complete the clustering of these massive images is particularly important. Conducting clustering analysis on scene images can extract scene semantic information from scene images, which is of great significance for scene recognition and retrieval. The data obtained by performing multiple feature extractions (such as text, color, shape, texture, etc.) on scene images is called multi-view scene image data, which can express the visual information of scene images more comprehensively and accurately. By utilizing the consistency and complementarity between different perspective information, better clustering effects can be obtained. However, due to factors such as shooting angle, light, and occluders, data of some perspectives is missing, resulting in the incomplete problem of multi-view scene images. Existing multi-view image clustering methods have good effects in dealing with the clustering problem of complete perspective data, but cannot be applied to the image clustering problem with incomplete perspectives, which to a certain extent increases the difficulty of multi-view scene image clustering. Therefore, more effective and robust technical methods are needed to solve the application of incomplete multi-view scene image clustering. Summary of the Invention
[0003] To solve the above problems, the present invention provides an incomplete multi-view scene image clustering method based on perspective exclusive pseudo-label propagation, and the method includes the steps:
[0004] Extract a complete sample set and an existing sample set from an incomplete multi-view scene image dataset. Solve the multi-view common subspace optimization problem for the features of the complete samples to obtain the perspective common subspace; perform spectral clustering on the perspective common subspace to obtain the pseudo-labels of the complete samples.
[0005] For each perspective, construct a perspective-specific semi-supervised classifier based on a graph convolutional neural network and randomly initialize the corresponding network weights; and construct a k-nearest neighbor graph matrix according to the existing samples of this perspective.
[0006] For each perspective, input the features of the existing samples and the corresponding k-nearest neighbor graph into the perspective-specific semi-supervised classifier to predict the pseudo-labels of the complete samples, and its loss function is a masked cross-entropy classification loss function.
[0007] For each perspective, use the stochastic gradient descent algorithm to minimize the masked classification loss function, and train the perspective-specific semi-supervised classifier until convergence. Obtain the cluster distribution on the existing samples from the converged classifier.
[0008] Perform a weighted average of the cluster distributions of the existing samples for all perspectives to obtain the cluster distribution of all samples; select the cluster with the highest probability for each sample as the cluster to which the sample belongs to obtain the final clustering result. Calculate the clustering accuracy on the incomplete multi-perspective scenario image dataset according to the obtained clustering result.
[0009] Furthermore, extract the complete sample set and the existing sample set from the dataset. The definitions of the complete sample set and the existing sample set are as follows:
[0010]
[0011]
[0012] where is the complete sample set, and ε v (v = 1, 2, …, m) is the existing sample set for each perspective. M ∈ {0, 1} n×m is the missing indicator matrix. If the j-th perspective of the i-th sample exists, then M ij = 1, otherwise M ij = 0; n and m represent the number of samples and the number of perspectives respectively; the feature of the complete sample is denoted as the feature of the existing sample is denoted as
[0013] Solve the multi-perspective common subspace optimization problem for the feature of the complete sample, and its definition is as follows:
[0014]
[0015] where X v is the feature of the complete sample, is the perspective-specific subspace matrix, is the perspective common subspace matrix. Perform spectral clustering on S to obtain the pseudo-labels of the complete samples, where n p is the number of complete samples, and n c is the number of clusters.
[0016] Furthermore, for each perspective v, construct a perspective-specific semi-supervised classifier based on a graph convolutional neural network and randomly initialize the network weights. The perspective-specific semi-supervised classifier consists of two graph convolutional modules, and its definition is as follows:
[0017]
[0018]
[0019] Among them, relu is a non-linear activation function layer, dropout is a random dropout layer, conv is a graph convolutional layer, and softmax is a multi-classification activation function layer; is the existing sample feature matrix, A v is the k-nearest neighbor graph matrix. is the output of the first graph convolutional module, is the cluster distribution of the existing samples of perspective v; n v are the existing samples of perspective v, d is the hidden layer size of the classifier, n c is the number of clusters.
[0020] For each perspective of the existing samples, construct a k-nearest neighbor graph matrix respectively, which is defined as follows:
[0021]
[0022] Among them, is the k-nearest neighbor graph matrix is the element in the i-th row and j-th column of, and i and j respectively represent two nodes on the k-nearest neighbor graph; is the k-neighborhood of node i, which consists of the k nodes closest to node i; if there is an edge between node i and node j, then otherwise
[0023] Furthermore, for each perspective, input the features of the existing samples and the corresponding k-nearest neighbor graph matrix into the perspective-specific semi-supervised classifier to predict the pseudo-labels of the complete samples, and its loss function is a masked cross-entropy classification loss function, which is defined as follows:
[0024]
[0025] Among them, is the loss function of perspective v, is the pseudo-label of the complete sample, is the cluster distribution predicted by the network; is the set of complete samples, represents taking out the part of the cluster distribution predicted by the network corresponding to the set of complete samples; n p is the number of complete samples, n c is the number of clusters.
[0026] Furthermore, for each perspective, use the stochastic gradient descent algorithm to minimize the masked classification loss function, and train the perspective-specific semi-supervised classifier until convergence. The update rule of the network weights is as follows:
[0027]
[0028] where Θ v is the weight of the semi-supervised classifier, α is the learning rate, is the loss function with respect to the network weight Θ v partial derivative. Repeatedly apply the above update rule until the value of the loss function converges. Then, based on the pseudo-labels of the complete samples and the cluster distribution predicted by the classifier, obtain the cluster distribution on the existing samples as follows:
[0029]
[0030] where is the cluster distribution of the existing sample i, y i is the pseudo-label of the complete sample i, is the cluster distribution predicted by the classifier for the existing sample i. The above formula shows that if the existing sample i is a complete sample, its cluster distribution is determined by the pseudo-label, otherwise it is determined by the cluster distribution predicted by the classifier.
[0031] Furthermore, perform a weighted average of the cluster distributions of the existing samples from all perspectives to obtain the cluster distribution of all samples, as follows:
[0032]
[0033] where M :,v represents the v-th column of the missing indicator matrix, represents the cluster distribution of all samples, Y v represents the cluster distribution on the existing samples. Select the cluster with the highest probability for each sample as the cluster to which the sample belongs to obtain the final clustering result, as follows:
[0034]
[0035] where c i represents the cluster label of sample i, represents the probability that sample i belongs to cluster j. According to the obtained clustering result c ∈ {1, 2,..., n c} n , calculate the clustering accuracy on the incomplete multi-view scenario image dataset.
[0036] The present invention provides a method for clustering incomplete multi-view scenario images based on perspective-exclusive pseudo-label propagation, having the following advantages:
[0037] (1) The method adopts a multi-view clustering framework, which can discover the potential connections between various visual features contained in the scene images, make full use of the consistency and complementarity among the multi-view scene image information, and can effectively improve the clustering quality.
[0038] (2) The method adopts a mutually exclusive view technology. Each view can independently construct a k-nearest neighbor graph and propagate pseudo-labels within the view, and can be parallelized across multiple processors to accelerate training, with a shorter running time when dealing with large-scale multi-view data sets.
[0039] (3) The method adopts a pseudo-label technology. Using the "labels" generated by the model can expand the number of supervision signals for image clustering. Especially in the case where "true" labels are scarce, pseudo-labels can improve the clustering performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 is a flowchart of a non-complete multi-view scene image clustering method based on mutually exclusive view pseudo-label propagation provided by the present invention.
[0042] Figure 2 is a schematic diagram of a non-complete multi-view scene image data set Scene-15. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further details the present invention in conjunction with specific embodiments and with reference to the drawings. It should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0044] Exemplary Method
[0045] As Figure 1 , the present invention provides a non-complete multi-view scene image clustering method based on mutually exclusive view pseudo-label propagation. The steps of the method are as follows:
[0046] Step S110: First, extract a complete sample set and an existing sample set from the non-complete multi-view scene image data set, and solve the multi-view common subspace optimization problem for the features of the complete samples. Its definition is as follows:
[0047]
[0048] Among them, X v is the feature of the complete sample, is the perspective-specific subspace matrix, is the perspective common subspace matrix. Utilize self-expression to learn the subspace S of the perspective v , Utilize perspective consistency to learn the shared subspace S, use the sparse subspace clustering method to obtain the perspective common subspace S; apply spectral clustering to S to obtain the pseudo-labels of the complete samples
[0049] Step S120: For each perspective v, construct a perspective-specific semi-supervised classifier based on a graph convolutional neural network and randomly initialize the network weights. The perspective-specific semi-supervised classifier consists of two graph convolutional modules, which are defined as follows:
[0050]
[0051]
[0052] Among them, relu is the non-linear activation function layer, dropout is the random dropout layer, conv is the graph convolutional layer, and softmax is the multi-classification activation function layer; use the rectified linear unit as the activation function of the layer, and insert two dropout layers to alleviate overfitting.
[0053] Construct a k-nearest neighbor graph matrix for each perspective with existing samples, which is defined as follows:
[0054]
[0055] Among them, is the element in the i-th row and j-th column of the k-nearest neighbor graph matrix , and i and j respectively represent two nodes on the k-nearest neighbor graph; is the k-neighborhood of node i, which consists of the k nodes closest to node i; if there is an edge between node i and node j, then otherwise
[0056] Step S130: For each perspective, input the features of the existing samples and the corresponding k-nearest neighbor graph matrix into the perspective-specific semi-supervised classifier to predict the pseudo-labels of the complete samples, and its loss function is the masked cross-entropy classification loss function, which is defined as follows:
[0057]
[0058] Among them, is the loss function of perspective v, is the pseudo-label of the complete sample, is the cluster distribution predicted by the network; is the set of complete samples, denotes taking out the part of the cluster distribution predicted by the network corresponding to the set of complete samples; n p is the number of complete samples, n c is the number of clusters.
[0059] Step S140: For each perspective, use the stochastic gradient descent algorithm to minimize the masked classification loss function and train the perspective-specific semi-supervised classifier until convergence. The update rule for the network weights is as follows:
[0060]
[0061] where Θ v is the weight of the semi-supervised classifier, α is the learning rate, is the loss function with respect to the network weight Θ v partial derivative. Repeatedly apply the above update rule until the value of the loss function converges. Then, based on the pseudo-labels of the complete samples and the cluster distribution predicted by the classifier, obtain the cluster distribution on the existing samples as follows:
[0062]
[0063] where, is the cluster distribution of existing sample i, y i is the pseudo-label of complete sample i, is the cluster distribution predicted by the classifier for existing sample i.
[0064] Step S150: Perform a weighted average on the cluster distributions of the existing samples for all perspectives to obtain the cluster distribution of all samples, as follows:
[0065]
[0066] where M :,v represents the v-th column of the missing indicator matrix, represents the cluster distribution of all samples, Y v represents the cluster distribution on the existing samples. Select the cluster with the highest probability for each sample as the cluster to which the sample belongs to obtain the final clustering result, as follows:
[0067]
[0068] Among them, c i represents the cluster label of sample i, and represents the probability that sample i belongs to cluster j. According to the obtained clustering result c ∈ {1, 2,..., n c} n .
[0069] Step S160: Calculate the clustering accuracy on the incomplete multi-view scene image dataset according to the obtained clustering result.
[0070] Through this embodiment, first, a complete sample set and an existing sample set are extracted from an incomplete multi-view scene image dataset. Pseudo-labels are generated from the complete samples, and then a perspective-specific semi-supervised classifier is constructed and a k-nearest neighbor graph for each perspective is constructed. Then, the perspective-specific semi-supervised classifier is trained until convergence. Finally, the cluster distribution of the existing samples obtained by the classifier is weighted and averaged to obtain the final result. After obtaining the clustering result, the clustering accuracy on the incomplete multi-view scene image dataset is calculated.
[0071] Furthermore, it is assumed that an incomplete multi-view scene image dataset is clustered according to this embodiment, and a clustering result with an accuracy higher than that of most methods will be obtained.
[0072] Specific implementation results
[0073] This embodiment uses a publicly available multi-view scene image dataset and simulates an incomplete multi-view scene image dataset according to different sample missing ratios. The details of the dataset are described as follows:
[0074] Scene-15 is a grayscale image dataset containing 15 indoor and outdoor scene categories, including scenes such as office, bedroom, forest, and street. This dataset has the following three perspectives:
[0075] Perspective 1 is the PHOG feature, also known as the hierarchical histogram of oriented gradients. The characteristic of this perspective is that the image is first pyramidally layered and then the local appearance and shape features of the image are characterized on each layer, representing the local contour features of the image.
[0076] Perspective 2 is the GIST feature. The characteristic of this perspective is to extract the global features of the image. After filtering the entire image, the means in each direction and at each scale are taken in the local area to extract the contour information of the image.
[0077] To verify the superiority of this embodiment (Ours), this embodiment is compared with several existing incomplete multi-view scene clustering methods, including methods such as BMVC, BSV, OPIMC, and FMVAC. The clustering accuracy (ACC) of these methods for five incomplete sample ratios on the above-mentioned public dataset will be compared. The specific data comparison is shown in Table 1.
[0078] Table 1 Clustering accuracy (%) of Scene-15 dataset
[0079]
[0080] From the data comparison in the above table, it can be clearly seen that Ours achieves the best performance and significantly improves the clustering accuracy of incomplete multi-view scene images. The quantitative results fully illustrate the superiority of Ours because Ours can better capture the view compatibility information and view complementarity information in incomplete multi-view scene images. Ours uses pseudo-label propagation to obtain the pseudo-labels of complete samples and constructs view-specific semi-supervised classifiers, thereby improving the clustering effect of incomplete multi-view scene images. A large number of experiments show that this method is superior to existing methods. Regarding the parameter settings of this embodiment, in all experiments, the parameter k of the k-nearest neighbor graph matrix is set to 5.
[0081] This embodiment proposes an incomplete multi-view scene image clustering method based on view-exclusive pseudo-label propagation, which is used to perform clustering analysis on common incomplete multi-view scene images in life. First, a complete sample set and an existing sample set are extracted from an incomplete multi-view scene image dataset to obtain the pseudo-labels of complete samples. Then, view-specific semi-supervised classifiers and k-nearest neighbor graph matrices are constructed, and the classifiers are trained until convergence to obtain the cluster distribution on the existing samples. The cluster distributions of the existing samples in all views are weighted and averaged to obtain the cluster distribution of all samples, and the final clustering result is obtained. The experimental results on five incomplete sample ratios of the Scene-15 scene image dataset show that this embodiment is more stable than other methods, less affected by the data missing rate, and has better clustering performance.
[0082] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principle of the present invention, and do not constitute a limitation to the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A non-complete multi-view scene image clustering method based on perspective-exclusive pseudo-label propagation, characterized in that The method includes the steps: S110. Extract a complete sample set and an existing sample set from an incomplete multi-view scene image dataset; solve the multi-view common subspace optimization problem for the features of the complete samples to obtain the view common subspace; perform spectral clustering on the view common subspace to obtain the pseudo-labels of the complete samples; S120. For each view, construct a view-specific semi-supervised classifier based on a graph convolutional neural network and randomly initialize the corresponding network weights; and construct a k-nearest neighbor graph matrix according to the existing samples of this view; S130. For each view, input the features of the existing samples and the corresponding k-nearest neighbor graph into the view-specific semi-supervised classifier to predict the pseudo-labels of the complete samples, and its loss function is the masked cross-entropy classification loss function; S140. For each view, use the stochastic gradient descent algorithm to minimize the masked classification loss function, and train the view-specific semi-supervised classifier until convergence; obtain the cluster distribution on the existing samples from the converged classifier; S150. Perform weighted averaging on the cluster distributions of the existing samples of all views to obtain the cluster distribution of all samples; Select the cluster with the highest probability for each sample as the cluster to which the sample belongs to obtain the final clustering result; calculate the clustering accuracy on the incomplete multi-view scene image dataset according to the obtained clustering result; In the above S110, extract the complete sample set and the existing sample set from the dataset; the definitions of the complete sample set and the existing sample set are as follows: Among them, is a complete sample set, and ε v (v = 1, 2, …, m) is the existing sample set of each perspective; M ∈ {0, 1} n×m is a missing indication matrix. If the v-th perspective of the i-th sample exists, then M iv = 1, otherwise M iv = 0; n and m respectively represent the number of samples and the number of perspectives; the feature of the complete sample is denoted as The feature of the existing sample is denoted as n p is the number of complete samples, and n v is the existing sample of perspective v; Features of complete samples Solve the multi-view common subspace optimization problem, which is defined as follows: Among them, X v is the feature of the complete sample, is the view-specific subspace matrix, is the view-common subspace matrix; perform spectral clustering on S to obtain the pseudo-labels of the complete samples n p is the number of complete samples, n c is the number of clusters.
2. The method for clustering incomplete multi-view scene images based on perspective-exclusive pseudo-label propagation according to claim 1, characterized in that For each view v, construct a view-specific semi-supervised classifier based on a graph convolutional neural network and randomly initialize the network weights. The view-specific semi-supervised classifier consists of two graph convolutional modules, and its definition is as follows: Among them, relu is a non-linear activation function layer, dropout is a random dropout layer, conv is a graph convolutional layer, and softmax is a multi-classification activation function layer; is the sample feature matrix, A v is the k-nearest neighbor graph matrix, is the output of the first graph convolutional module, is the cluster distribution of samples existing in perspective v; n v are the samples existing in perspective v, d is the hidden layer size of the classifier, n c is the number of clusters; Construct a k-nearest neighbor graph matrix for each view of the existing samples respectively, and the definition is as follows: Among them, is the element in the \(e\)-th row and \(f\)-th column of the \(k\)-nearest neighbor graph matrix , where \(e\) and \(f\) respectively represent two nodes on the \(k\)-nearest neighbor graph; is the \(k\)-neighborhood of node \(e\), which consists of the \(k\) nodes closest to node \(e\); if there is an edge between node \(e\) and node \(f\), then otherwise 3. The incomplete multi-view scene image clustering method based on perspective-exclusive pseudo-label propagation according to claim 1, wherein For each view, input the features of the existing samples and the corresponding k-nearest neighbor graph matrix into the view-specific semi-supervised classifier to predict the pseudo-labels of the complete samples, and its loss function is the masked cross-entropy classification loss function, and the definition is as follows: Among them, is the loss function of the viewing angle v, is the pseudo-label of the complete sample, is the cluster distribution predicted by the network; is the set of complete samples, means to take out the cluster distribution predicted by the network corresponding to the set of complete samples; n p is the number of complete samples, n c is the number of clusters.
4. The non-complete multi-view scene image clustering method based on perspective-exclusive pseudo-label propagation according to claim 1, characterized in that For each view, use the stochastic gradient descent algorithm to minimize the masked classification loss function, and train the view-specific semi-supervised classifier until convergence. The update rule of the network weights is as follows: where Θ v is the weight of the semi - supervised classifier, α is the learning rate, is the loss function with respect to the network weight Θ v partial derivative; Repeatedly apply the above update rule until the value of the loss function converges; Then, based on the pseudo - labels of the complete samples and the class cluster distribution predicted by the classifier, obtain the class cluster distribution on the existence samples as follows: wherein, is the cluster distribution of the existing sample i, is the pseudo-label of the complete sample i, is the classifier predicted cluster distribution of the existing sample i; the above formula indicates that if the existing sample i is a complete sample, its cluster distribution is determined by the pseudo-label, otherwise it is determined by the cluster distribution predicted by the classifier.
5. The non-complete multi-view scene image clustering method based on perspective-exclusive pseudo-label propagation according to claim 1, characterized in that, Perform weighted averaging on the cluster distributions of the existing samples of all views to obtain the cluster distribution of all samples, as follows: Among them, M :,v represents the v-th column of the missing indication matrix, represents the cluster distribution of all samples, represents the cluster distribution on existing samples; for each sample, the cluster with the highest probability is selected as the cluster to which the sample belongs, and the final clustering result is obtained as follows: Among them, c i represents the cluster label of sample i, represents the probability that sample i belongs to cluster j; According to the obtained clustering result c ∈ {1, 2,..., n c} n , calculate the clustering accuracy on the incomplete multi-view scene image dataset.