Image Clustering Method Based on Image-Text Pretrained Model
By using WordNet nouns to construct text representation of images in the pre-trained graphic and text model, and using the method of mutual distillation of graphic and text modality, the problem of poor image clustering effect under unknown category names is solved, and efficient image clustering effect is achieved.
Patent Information
- Application Number
- CN202311112435.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-08-30
AI Technical Summary
In the case of unknown category names, it is difficult for the prior art to effectively utilize the text mode of the pre-trained graphics model, resulting in poor image clustering effect and large computational overhead.
By obtaining all nouns in WordNet, using the graphic and text pre-training model to select candidate words for text mode, construct their representation in text mode for each image, and obtain image clustering results using the mutual distillation of graphic and text modes.
In the case of unknown category names, the text mode of the pre-trained graphics and text models is effectively utilized, which improves the performance and efficiency of image clustering without additional model training and tuning.
Smart Images

Figure CN117173441B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image clustering method based on a text-image pre-training model. Background Art
[0002] As one of the classic tasks in machine learning, image clustering has a development history of several decades. Early image clustering work mainly focused on designing clustering strategies to discover data clusters from given raw image data. The most classic one is the proposed k-means clustering method. Although early clustering methods achieved good results on some simple data, they were not satisfactory when faced with linearly inseparable image data with higher dimensions and complexities in the real world. To better handle complex data, deep clustering methods were proposed, which utilize the powerful feature extraction ability of neural networks to first extract more discriminative low-dimensional representations from the raw data, thus greatly improving the image clustering effect. For example, an image clustering method based on autoencoders optimizes a deep neural network by designing a clustering objective, and simultaneously achieves feature extraction and data clustering. It can be seen that the quality of image clustering methods is closely related to the quality of data representations. In recent years, the contrastive learning paradigm has made great progress in unsupervised representation learning. Without knowing the class labels of each image, it constructs self-supervised signals through data augmentation to help the model extract compact low-dimensional representations. Thanks to the success of contrastive learning, a series of image clustering methods based on contrastive learning have been proposed in recent years, elevating the performance of image clustering to a new level. For example, contrastive learning is simultaneously performed in the row space and column space of the feature matrix to achieve image clustering for large-scale online data. Recently, unsupervised representation learning has gradually developed from a single image modality to an image-text multimodality. By learning from hundreds of millions of image-text data pairs on the Internet, text-image pre-training models have powerful representation learning capabilities and have achieved excellent performance in tasks such as image-text retrieval and image classification. However, there is currently no in-depth research on how to apply text-image pre-training models to image clustering.
[0003] Among them, the CLIP text-image pre-training model can classify images without knowing the category of each image. Specifically, given the category names of K categories to be classified (such as "Cat", "Dog", "Car", etc.), the CLIP text-image pre-training model first constructs a prompt such as "A photo of
CLASS
CLASS
[0004] Although the above solution can achieve image classification, its feasibility depends on the prior knowledge of the category names, that is, it is necessary to pre-give the category names to be classified such as "Cat", "Dog", "Car", etc. However, this prior information of the category names cannot be obtained in the unsupervised image clustering scenario. Therefore, this paradigm cannot achieve image clustering. Therefore, in the image clustering task with unknown category names, a directly feasible image clustering method based on the CLIP image-text pre-training model is to use its pre-trained image encoder to extract image features and further use the traditional k-means method to achieve image clustering. However, this solution fails to utilize the text modality with compact semantic information, resulting in limited clustering performance. Summary of the Invention
[0005] Aiming at the above deficiencies in the prior art, an image clustering method based on an image-text pre-training model provided by the present invention solves the problems that the text modality of the image-text pre-training model cannot be effectively utilized in the case of unknown category names, resulting in poor image clustering effect and large computational overhead.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] This solution provides an image clustering method based on an image-text pre-training model, including the following steps:
[0008] S1. According to all nouns obtained from WordNet, use the image-text pre-training model to select candidate words in the text modality, retrieve the candidate words for each image, and construct its corresponding representation in the text modality;
[0009] S2. According to the image representation and its corresponding representation in the text modality, use the method of cross-modal mutual distillation to obtain the image clustering result.
[0010] The beneficial effects of the present invention are as follows: By dividing all nouns in WordNet into the semantic centers of images, selecting representative nouns for retrieval to construct the corresponding representations of each image in the text modality, and using the consistency of neighbor clustering assignments to coordinate the image and text modalities, the performance of image clustering is further improved by additionally training a clustering network on the image-text representations. The present invention can effectively utilize the text modality of the image-text pre-training model to mine semantic information, enhance the discriminability of image representations, and thus improve the image clustering effect without additional model training and model tuning in the case of unknown class names. The proposed paradigm should retain the advantage of low computational overhead as much as possible while improving performance, thereby enhancing the usability of the proposed solution.
[0011] Further, the step S1 includes the following steps:
[0012] S101. Respectively obtain the image encoder and text encoder of the pre-trained image-text pre-training model, all nouns in WordNet, and the images to be clustered;
[0013] S102. Input the images to be clustered into the image encoder to obtain image representations, and input all nouns in WordNet into the text encoder to obtain text representations;
[0014] S103. Calculate the image semantic centers according to the image representations using the k-means clustering algorithm;
[0015] S104. Divide all nouns into k image semantic centers according to the similarity between the text representations and the image semantic centers. The probability p(l|t i' ) of the i'-th noun belonging to the l-th clustering and the image semantic center is expressed as follows:
[0016]
[0017] where sim(·) represents cosine similarity, t i' represents the text representation obtained by the i'-th noun after passing through the text encoder, s l represents the l-th image semantic center, s j represents the j-th image semantic center, and k represents the number of image semantic centers;
[0018] S105. Retain the top 5 nouns with the highest probability for each image semantic center and use them as candidate words in the text modality;
[0019] S106. Construct the corresponding representation of each image in the text modality by retrieving the candidate words for each image.
[0020] The beneficial effect of the above further scheme is that the present invention aims to construct a text modality representation for each image by constructing a text modality to mine the semantic information of the text modality.
[0021] Furthermore, the expression of the image semantic center in step S103 is as follows:
[0022]
[0023] k=max{N / 300,K*3}
[0024] Among them, s l represents the lth image semantic center, v i represents the image representation obtained after the i-th image passes through the image encoder, Indicates v i Belongs to the lth cluster, is an indicator function if and only if v i It is 1 when it belongs to the lth cluster and 0 in other cases. N represents the number of sample points and K represents the number of target clusters.
[0025] The beneficial effect of the above further scheme is: estimating an appropriate number of semantic centers of an image, avoiding selecting too large a number of semantic centers, resulting in too fine a semantic granularity which is not conducive to clustering, or too small a number of semantic centers, resulting in an inability to accurately describe the semantic information of the image, especially the image located at the boundary of the cluster.
[0026] Furthermore, the expression of the representation corresponding to the text modality in step S106 is as follows:
[0027]
[0028]
[0029] in, represents the representation in the text modality corresponding to the i-th image, M represents the number of all candidate nouns retained after screening, represents the similarity between the j'th candidate noun and the i-th image, represents the j'th candidate noun constituting the text modality, v i represents the image representation obtained after the i-th image passes through the image encoder, Indicates the smoothness of the control retrieval. Represents the k'th candidate noun constituting the text modality.
[0030] The beneficial effect of the above further solution is to prevent different images from collapsing to the same feature point in the text space. Collapsing to the same feature point will cause the personalized information of the image to be completely lost and will directly affect the subsequent search for neighbors.
[0031] Further, the step S2 includes the following steps:
[0032] S201. Determine whether to not select an additional clustering network. If so, proceed to step S202; otherwise, select an additional clustering network and proceed to step S203;
[0033] S202. Concatenate the text representation and the image representation, and use the k-means clustering algorithm to obtain the image clustering result;
[0034] S203. Respectively construct a text clustering network and an image clustering network, find 50 nearest neighbors for each image representation in the image modality, and find 50 nearest neighbors for the text representation in the text modality;
[0035] S204. Obtain a batch of data composed of image representations and text representations, and randomly sample a neighbor image representation and a text representation from their neighbors;
[0036] S205. Input the sampled neighbor image representation and text representation into the image clustering network and the text clustering network respectively to obtain the clustering assignment of the image and the clustering assignment of the text. Among them, the text representation in steps S202 to S205 is the corresponding representation in the text modality obtained in step S1;
[0037] S206. Calculate the loss function L Overall , and use the loss function L Overall to optimize the image clustering network and the text clustering network;
[0038] S207. Determine whether the optimized image clustering network and text clustering network have converged. If so, proceed to step S208; otherwise, return to step S204;
[0039] S208. Input the image to be clustered into the optimized image clustering network, and select the prediction with the highest probability as the image clustering result.
[0040] The beneficial effect of the above further solution is that the present invention utilizes the mutual distillation of text and image modalities, aiming to synergistically combine the information of the text and image modalities to achieve better image clustering effects.
[0041] Further, the expression of the clustering assignment of the image in step S205 is as follows:
[0042]
[0043] Where P and P NBoth represent the image clustering assignment matrix of n*K, where n represents the batch size in the batch optimization of the deep neural network, K represents the number of target clusters, and p i and respectively represent the image clustering assignments of the i-th image and its neighbor, where i = 1, 2,..., n;
[0044] The expression of the clustering assignment of the text is as follows:
[0045]
[0046] where Q and Q N both represent the text clustering assignment matrix of n*K, and q x and respectively represent the text clustering assignments of the x-th text and its neighbor, where x = 1, 2,..., n.
[0047] Furthermore, the loss function L Overall in the step S206 has the following expression:
[0048] L Overall = L Dis + L Con - a·L Bal
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055] where L Overall represents the loss function, L Dis represents the image-text mutual distillation loss function, L Con represents the clustering confidence loss function, a represents the weight parameter, L Bal represents the weight parameter, represents the graph-to-text distillation loss function, represents the text-to-graph distillation loss function, i” represents the i”-th sample including the image / text pair, K and K' both represent the number of target clusters, sim(·) represents the cosine similarity, and respectively represent the clustering assignment matrices P, P N 、Q and Q Nthe i-th column of represents the temperature coefficient, n represents the batch size in batch optimization of the deep neural network, represents p i transpose of, p i represents the clustering assignment vector of the i-th image in the sample, 1 i=x is the indicator function, which is 1 if and only if i is equal to x, and 0 otherwise, q x represents the text clustering assignment of the x-th text, k' represents the k'-th column of the clustering assignment matrix, represents the overall distribution of all image clustering assignments in the current batch of data the i''-th element of represents the overall distribution of all text clustering assignments in the current batch of data the i''-th element of represents a space of dimension K, and respectively represent the clustering assignment matrices P N and Q N the k'-th column of.
[0056] The beneficial effect of the above further solution is: by fully synergizing the neighbor information of the image and text modalities, images with similar semantics can be divided into the same category, thereby improving the final clustering effect. Brief Description of the Drawings
[0057] Figure 1 is the flowchart of the method of the present invention. Detailed Embodiments
[0058] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0059] Embodiment
[0060] Before explaining the present invention, the following terms are first explained:
[0061] Image-Text Pretrained Model: A large-scale pre-trained model that learns the complex relationships and semantic representations between images and texts through joint pre-training on a large amount of paired image and text data. This model can extract meaningful features from images and texts for various downstream tasks including image-text retrieval, image segmentation, and semantic segmentation.
[0062] Image Clustering: A fundamental unsupervised machine learning method that aims to partition images with similar features or semantics into the same cluster and dissimilar images into different clusters without relying on data labels. Image clustering techniques can automatically discover patterns and structures in image data without using predefined class labels. Image clustering has wide applications in fields such as image retrieval, image library management, and unsupervised learning, and helps to organize and understand large-scale image data.
[0063] As Figure 1 shown, the present invention provides an image clustering method based on a text-image pre-trained model, and its implementation method is as follows:
[0064] S1. According to all nouns obtained from WordNet, use the text-image pre-trained model to select candidate words in the text modality, retrieve the candidate words for each image, and construct its corresponding representation in the text modality:
[0065] S101. Respectively obtain the image encoder and text encoder of the pre-trained text-image pre-trained model, all nouns in WordNet, and the images to be clustered;
[0066] S102. Input the images to be clustered into the image encoder to obtain image representations, and input all nouns in WordNet into the text encoder to obtain text representations;
[0067] S103. According to the image representations, use the k-means clustering algorithm to calculate the image semantic centers;
[0068] S104. Divide all nouns into k image semantic centers according to the similarity between the text representations and the image semantic centers;
[0069] S105. Retain the top 5 nouns with the highest probability for each image semantic center and use them as candidate words in the text modality;
[0070] S106. Construct its corresponding representation in the text modality by retrieving the candidate words for each image;
[0071] S2. According to the image representations and their corresponding representations in the text modality, use the text-image modality mutual distillation method to obtain the image clustering results, and its implementation method is as follows:
[0072] S201. Determine whether to select an additional clustering network. If so, proceed to step S202; otherwise, select an additional clustering network and proceed to step S203;
[0073] S202. Concatenate the text representations and the image representations, and use the k-means clustering algorithm to obtain the image clustering results;
[0074] S203. Construct a text clustering network and an image clustering network respectively, find 50 nearest neighbors for each image representation in the image modality, and find 50 nearest neighbors for the text representation in the text modality;
[0075] S204. Obtain a batch of data composed of image representations and text representations, and randomly sample a neighbor image representation and a text representation from their neighbors;
[0076] S205. Input the sampled neighbor image representation and text representation into the image clustering network and the text clustering network respectively to obtain the clustering assignment of the image and the clustering assignment of the text. Among them, the text representations in steps S202 to S205 are all the representations corresponding to the text modality obtained in step S1;
[0077] S206. Calculate the loss function L Overall , and use the loss function L Overall to optimize the image clustering network and the text clustering network;
[0078] S207. Determine whether the optimized image clustering network and text clustering network have converged. If so, go to step S208; otherwise, return to step S204;
[0079] S208. Input the image to be clustered into the optimized image clustering network, and select the prediction with the highest probability as the image clustering result.
[0080] In this embodiment, since class names are not available in the image clustering task, all nouns from WordNet are used as candidate words in the text modality. The key problem is how to select an appropriate set of nouns to form the text space. To accurately cover the image semantics and highly distinguish different categories of images, the present invention uses k-means to obtain the image semantics, and selects an appropriate k value in k-means according to the number of sample points N and the target number of clusters K. Specifically, the present invention calculates the k value through the following formula:
[0081] k = max{N / 300, K * 3}
[0082] After that, the present invention uses k-means to calculate the semantic center of the image, specifically as follows:
[0083]
[0084] where s l represents the l-th image semantic center, v i represents the image representation obtained by the i-th image through the image encoder, represents that v i belongs to the l-th cluster, is an indicator function that is 1 if and only if v i belongs to the l-th cluster, and 0 otherwise. N represents the number of sample points, and K represents the number of target clusters.
[0085] In this embodiment, after obtaining the image semantic center, in order to select a representative noun set, all nouns in WordNet are divided into k' image semantic centers. The probability p(l|t i' ) of the i'-th noun belonging to the l-th cluster and the image semantic center is expressed as follows:
[0086]
[0087] where sim(·) represents cosine similarity, t i' represents the text representation obtained by the text encoder for the i'-th noun, s l represents the l-th image semantic center, s j represents the j-th image semantic center, and k represents the number of image semantic centers.
[0088] In this embodiment, the 5 nouns with the highest corresponding probability for each semantic center are retained as candidate words for constructing the text space. After selecting the representative noun set, the present invention constructs the text modality by retrieving the most relevant noun for each image. Specifically:
[0089]
[0090]
[0091] where, represents the representation in the text modality corresponding to the i-th image, M represents the number of all candidate nouns retained after screening, represents the similarity between the j'-th candidate noun and the i-th image, represents the j'-th candidate noun that makes up the text space, v i represents the image representation obtained by the image encoder for the i-th image, represents the smoothness degree for controlling the retrieval, with the default setting of 0.005, represents the k'-th candidate noun that makes up the text modality.
[0092] So far, the representation of each image in the text modality has been constructed. At this time, the similarity between the text and the image can be calculated through the concatenated representation Directly applying the traditional k-means clustering method on it to achieve image clustering. Since the compact semantics from the text modality are incorporated, the concatenated representation has better discriminability, thus obtaining better image clustering results compared to directly using k-means on the image representation. It should be noted that the construction process of the above text modality does not require any additional training and model tuning, and the computational overhead of the noun selection and retrieval process is almost negligible.
[0093] In this embodiment, although directly concatenating the text and image representations and then using k-means can significantly improve the image clustering effect, this paradigm is still relatively simple and fails to fully synergize the two modalities of text and image. Therefore, the present invention further proposes a method for mutual distillation between the text and image modalities to further improve the clustering performance by training an additional clustering network.
[0094] Specifically, the present invention finds 50 most similar images for each image in the image modality to form a neighbor set N(v i ), and introduces a clustering network f to make clustering assignments for each image representation. In each iteration, calculate the clustering assignment of all images and a randomly selected image in their neighbor sets, denoted as:
[0095]
[0096] where, P and P N both represent the n*K image clustering assignment matrix, n represents the batch size in the batch optimization of the deep neural network, K represents the number of target clusters, p i and respectively represent the image clustering assignments of the i-th image and its neighbor, where i = 1, 2,..., n.
[0097] Similarly, introduce another clustering network g to make clustering assignments for each text representation. Also find 50 most similar texts for each text representation to form a neighbor set In each iteration, calculate the clustering assignment of all texts and a randomly selected text in their neighbor sets, denoted as:
[0098]
[0099] where, Q and Q N both represent the n*K text clustering assignment matrix, q x and respectively represent the text clustering assignments of the x-th text and its neighbor, where x = 1, 2,..., n.
[0100] To synergize the two modalities of images and texts, the present invention requires that the clustering network has similar clustering assignments for an image and its neighboring text modalities, and also has similar clustering assignments for a text and its neighboring image modalities. To achieve this goal, the following loss function is designed.
[0101]
[0102]
[0103]
[0104] On the one hand, this loss function can achieve the synergy of image and text modalities through the consistency of clustering assignments among cross-modal neighbors. On the other hand, it can expand the differences between different clusters.
[0105] In addition, to make the training process more stable, two other regular-term loss functions are designed. First, to encourage the model to make more confident clustering assignments, the following loss function is proposed:
[0106]
[0107] This loss function is minimized when both p i and q i are in one-hot encoding, so it can improve the confidence of clustering assignments. Additionally, to prevent the model from assigning a large number of images and texts to individual clusters, the following loss function is proposed.
[0108]
[0109]
[0110] Combining the above three loss functions, the following loss function is used to optimize the clustering networks f and g for image and text modalities:
[0111] L Overall = L Dis + L Con - a·L Bal
[0112] where L Overall represents the loss function, L Dis represents the image-text mutual distillation loss function, L Con represents the clustering confidence loss function, a represents the weight parameter, L Bal represents the weight parameter, represents the graph-to-text distillation loss function, Denote the text-to-image distillation loss function. "i" represents the i-th sample including the image / text pair. Both K and K' represent the number of target clusters. sim(·) represents cosine similarity. and represent the i-th columns of the clustering assignment matrices P, P N , Q, and Q N respectively. Denote the temperature coefficient. n represents the batch size in the batch optimization of the deep neural network. Denote the transpose of p i . p i represents the clustering assignment vector of the i-th image in the sample. 1 i=x is the indicator function, which is 1 if and only if i equals x, and 0 otherwise. q x represents the text clustering assignment of the x-th text. k' represents the k'-th column of the clustering assignment matrix. Denote the i'''-th element of the overall distribution of all image clustering assignments in the current batch of data . Denote the i'''-th element of the overall distribution of all text clustering assignments in the current batch of data . Denote the space of dimension K. and represent the k'-th columns of the clustering assignment matrices P N and Q N respectively.
[0113] It should be noted that the above loss function is only used to optimize the additionally introduced clustering network and does not modify the pre-trained text encoder and image encoder of CLIP. Therefore, its overall training cost is small. In the experiments of the present invention, the present invention only needs 1 minute to train on 60,000 images and has a wide range of application scenarios. Both the clustering networks f and g are 3-layer fully connected neural networks, where the dimensions of the input layer and the hidden layer are equal to the dimensions of the image and text features, and the dimension of the output layer is equal to the number of target clusters. The present invention uses gradient descent to optimize the network until convergence. After training is completed, just input the image to be clustered into the clustering network f to obtain its clustering assignment, thus realizing image clustering.
[0114] The present invention improves the performance of image clustering by utilizing the semantic information of the text modality in the image-text pre-training model. In the experiment, the more advanced methods currently used in the world are compared, including K-means clustering method, spectral clustering method (SC, NMF), hierarchical clustering method (AC, JULE), autoencoder method (AE, DAE, DeCNN, VAE), generative adversarial network method (DCGAN), deep clustering method (DEC, DCCM, PICA, CC) and other advanced methods, as well as the baseline method of using k-means on the image features extracted by CLIP, and experimental comparisons are carried out on the object image datasets CIFAR-10 and ImageNet-10. The commonly used indicator for measuring clustering effect, namely normalized mutual information (NMI), is used as a quantitative indicator of the experiment to verify the effect of the algorithm. The NMI value range is 0 to 1. The larger the number, the better the effect. When it is 1, it means that the algorithm can completely and correctly cluster the data correctly. The NMI calculation method is as follows:
[0115]
[0116] Among them, Y represents the category information predicted by the algorithm, C represents the actual category information of the data, H(·) represents information entropy, and I(Y; C) represents mutual information.
[0117] Experiment 1: Using the CIFAR-10 dataset, which contains 60,000 images from 10 object categories, the experimental data category information and sample quantity distribution are shown in Table 1:
[0118] Table 1
[0119]
[0120] The experimental results are shown in Table 2:
[0121] Table 2
[0122]
[0123]
[0124] As can be seen from the table, compared with other clustering methods, this method has a significant improvement in the normalized mutual information indicator, which means that it can correctly cluster object image data in practical applications and avoid wasting a lot of human resources on image classification.
[0125] Experiment 2: Use the ImageNet-10 dataset, which is a subset of the large image dataset ImageNet. It contains 13,000 images from 10 object categories. The experimental data category information and sample quantity distribution are shown in Table 3:
[0126] Table 3
[0127] Penguin Dog Leopard Airplane Airship Ship Football Sedan Truck Orange 1300 1300 1300 1300 1300 1300 1300 1300 1300 1300
[0128] The experimental results are shown in Table 4:
[0129] Table 4
[0130]
[0131]
[0132] As can be seen from the table, this method has a relatively large improvement in the normalized mutual information index compared with other clustering methods, which means that it can correctly cluster object image data well in practical applications, avoiding the consumption of a large amount of human resources for image classification.
[0133] The present invention has been verified on the CLIP text-image pre-training model, but the proposed image clustering paradigm based on the text-image pre-training model is also applicable to other text-image pre-training models. Similarly, the present invention has no requirements for the selection of text and image encoders and can be applicable to various encoders such as Transformer, ResNet, and ViT. In addition, k-means is used to calculate the image semantic center in the process of constructing the text modality, and the same function can also be achieved by other clustering methods in this step.
Claims
1. An image clustering method based on a text-image pre-trained model, characterized in that, it includes the following steps: S1. According to all nouns obtained from WordNet, use the text-image pre-trained model to select candidate words in the text modality, retrieve the candidate words for each image, and construct its corresponding representation in the text modality; The step S1 includes the following steps: S101. Respectively obtain the image encoder and text encoder of the pre-trained text-image pre-trained model, all nouns in WordNet, and the images to be clustered; S102. Input the images to be clustered into the image encoder to obtain image representations, and input all nouns in WordNet into the text encoder to obtain text representations; S103. According to the image representations, use the k-means clustering algorithm to calculate the image semantic centers; S104. Divide all nouns into k image semantic centers according to the similarity between the text representation and the image semantic center. Among them, the probability that the l th noun belongs to the th clustering and the image semantic center is expressed as follows: Among them, sim ( . ) represents the cosine similarity, represents the text representation obtained by passing the -th noun through the text encoder, represents the l -th image semantic center, represents the j -th image semantic center, k represents the number of image semantic centers; S105. Retain the top 5 nouns with the highest probability for each image semantic center and use them as candidate words in the text modality; S106. By retrieving the candidate words for each image, construct its corresponding representation in the text modality; S2. According to the image representations and their corresponding representations in the text modality, use the text-image modality mutual distillation method to obtain the image clustering results.
2. The image clustering method based on the text-image pre-trained model according to claim 1, characterized in that, the expression of the image semantic center in the step S103 is as follows: in, Indicates l Image semantic center, Indicates i The image representation obtained after the image encoder passes through the image. express Belong to l clusters, is an indicator function if and only if Belong to l It is 1 when there are 1 clusters, and 0 in other cases. N represents the number of sample points, K Indicates the number of target clusters.
3. The image clustering method based on the text-image pre-trained model according to claim 1, characterized in that, the expression of the corresponding representation in the text modality in the step S106 is as follows: Among them, represents the representation in the text modality corresponding to the i th image, M represents the number of all candidate nouns retained after screening, represents the th candidate noun and the i th image similarity, represents the th candidate noun that makes up the text modality, represents the i th image representation obtained after the image encoder processes the th image, represents the control of the smoothness of the retrieval, represents the th candidate noun that makes up the text modality.
4. The image clustering method based on the text-image pre-trained model according to claim 1, characterized in that, the step S2 includes the following steps: S201. Determine whether to not select an additional clustering network. If so, go to step S202; otherwise, select an additional clustering network and go to step S203; S202. Concatenate the text representations and image representations and use the k-means clustering algorithm to obtain the image clustering results; S203. Respectively construct a text clustering network and an image clustering network, find 50 nearest neighbors for each image representation in the image modality, and find 50 nearest neighbors for the text representation in the text modality; S204. Obtain a batch of data composed of image representations and text representations, and randomly sample a neighbor image representation and a text representation from their neighbors; S205. Input the sampled neighbor image representation and text representation into the image clustering network and the text clustering network respectively to obtain the clustering assignment of the image and the clustering assignment of the text. Among them, the text representations in steps S202 to S205 are all the corresponding representations in the text modality obtained in step S1; S206. Calculate the loss function according to the clustering assignment of the image and the clustering assignment of the text , and use the loss function to optimize the image clustering network and the text clustering network; S207. Determine whether the optimized image clustering network and text clustering network have converged. If so, go to step S208; otherwise, return to step S204; S208. Input the images to be clustered into the optimized image clustering network and select the prediction with the highest probability as the image clustering result.
5. The image clustering method based on the text-image pre-trained model according to claim 4, characterized in that, The expression for the clustering assignment of the image in step S205 is as follows: Among them, and both represent the image clustering assignment matrix of n represents the batch size in the batch optimization of the deep neural network, K represents the number of target clusters, and respectively represent the image clustering assignments of the i th image and its neighbor, where ; The expression for the clustering assignment of the text is as follows: Among them, and both represent 's text clustering assignment matrix, and respectively represent the text clustering assignments of the x th text and its neighbors, .
6. The image clustering method based on the text and image pre-training model according to claim 4, characterized in that The loss function in step S206 has the following expression: Among them, represents the loss function, represents the image-text mutual distillation loss function, represents the clustering confidence loss function, represents the weight parameter, represents the weight parameter, represents the graph-to-text distillation loss function, represents the text-to-graph distillation loss function, represents the th sample including the image / text pair, K and both represent the number of target clusters, sim ( . ) represents the cosine similarity, 、 、 and respectively represent the 、 、 and th column of the clustering assignment matrix , represents the temperature coefficient, n represents the batch size in the batch optimization of the deep neural network, represents transpose of, represents the clustering assignment vector of the i th image in the sample, is the indicator function, which is 1 if and only if i is equal to x , and 0 otherwise, represents the text clustering assignment of the x th text, represents the th column of the clustering assignment matrix, represents the overall distribution of all image clustering assignments in the current batch of data th element of, represents the overall distribution of all text clustering assignments in the current batch of data th element of, represents a space of dimension , and respectively represent the and th column of the clustering assignment matrix .