Image clustering method and system based on text guidance
Through the text-guided image clustering method, text feature screening and weighting processing, combined with sample replacement and pre-trained network model training, the shortcomings of traditional clustering methods in high-dimensional data and complex nonlinear relationships are solved, and efficient and accurate image clustering is achieved.
Patent Information
- Application Number
- CN202510305819.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional clustering methods are difficult to deal with high-dimensional data and complex nonlinear relationships, resulting in insufficient accuracy and reliability of image clustering results, especially in large-scale label-free image datasets.
By filtering and weighting according to the similarity between each image clustering center and text feature, the text features of each image are obtained, and the image clustering model is obtained through sample replacement and pre-training network model training, which improves the accuracy and clustering efficiency of text and image matching.
It improves the accuracy and efficiency of image clustering, is suitable for clustering of large-scale label-free image datasets, and improves the accuracy and stability of clustering results.
Smart Images

Figure CN120234432A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image clustering, and particularly relates to a text-guided image clustering method and system. Background Art
[0002] With the wide application of clustering methods in fields such as image segmentation and object detection, text data analysis, market segmentation, and bioinformatics, traditional clustering methods have limitations in dealing with modern complex data. Taking the classic K-Means algorithm as an example, based on a simple distance metric, it is difficult to handle high-dimensional data and complex non-linear relationships, and it is difficult to deeply explore the potential structure of data, resulting in insufficient accuracy and reliability of clustering results.
[0003] Currently, when performing image clustering, matching text and images to achieve guided image clustering can solve the problems existing in traditional clustering methods. However, in current methods, due to the lack of reasonable text screening and feature data update methods, the accuracy and efficiency of the clustering model in matching text and images are not high, and it is not suitable for clustering large-scale unlabeled image datasets. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a text-guided image clustering method and system. The present invention screens text features according to the similarity between each image clustering center and text features, obtains the text features of each image through a weighted method for the screened text features, and performs sample replacement according to a preset similarity range. According to the obtained image data and the replaced samples, a preset network model is trained to obtain an image clustering model, which improves the accuracy of text and image matching and the clustering efficiency, and is suitable for clustering large-scale unlabeled image datasets.
[0005] To achieve the above object, the present invention is implemented by the following technical solutions:
[0006] In a first aspect, the present invention provides a text-guided image clustering method, including:
[0007] Obtain image data;
[0008] According to the obtained image data and the image clustering model, obtain an image clustering result;
[0009] Among them, when training the image clustering model, text descriptions of the image data in the training set are generated, features of the images and texts are extracted respectively, text features are screened according to the similarity between each image clustering center and the text features, and the text features after screening are weighted to obtain the text features of each image; according to a preset similarity range, sample replacement is performed, and the preset network model is trained according to the obtained image data and the replaced samples to obtain the image clustering model.
[0010] Furthermore, perform word segmentation on the text description, extract the feature information of the image data and the text after word segmentation to obtain image features and text feature information; screen the text feature information to obtain the text features of each image; randomly select nearest neighbor samples from the image features and the screened text features, and delimit a range for each sample according to the similarity; input the extracted image features and the text features after sample selection into a preset mapping head, adjust the dimensions and optimize the network through instance-level contrast loss; input the extracted image features and text features, as well as the selected text features into a clustering head, adjust the dimensions to match the number of clusters, and train the network with cluster-level contrast loss.
[0011] Furthermore, calculate the similarity between each image sample and a preset number of texts after screening, use the similarity as a weight factor, and calculate the text embedding through weighted calculation.
[0012] Furthermore, the screened text index set R j is:
[0013]
[0014] where j is the cluster referred to; sim is the cosine similarity; t is the text modality; l is the subscript of the text embedding; is the l-th text embedding; c j is the j-th clustering center; μ is the set selection threshold; arg top-μ is to select the top μ text embeddings with the highest similarity to the clustering center c by calculating the similarity between the text embedding j and the clustering center c.
[0015] Furthermore, calculate the similarity matrix between samples in the same batch; for each sample, select a preset number of samples with the top-ranked similarities from the rows of the similarity matrix to form a candidate set, and randomly select a sample from the candidate set for feature replacement.
[0016] Furthermore, the instance-level contrast loss is:
[0017] f(z a , z b ) = sim(z a , zb )
[0018]
[0019] where z a is the input image embedding; z b is the input text embedding; N is the number of samples; b is the corresponding subscript when traversing all samples; τ1 is the instance-level temperature coefficient; and are the low-dimensional feature vectors of the a-th sample pair; is the instance-level multi-modal contrast loss from the image to the text direction.
[0020] Furthermore, the cluster-level contrast loss is:
[0021]
[0022] Jointly optimize L ins and L clu , and at the same time add the regularization loss L nom , the regularization loss is:
[0023]
[0024] where w b is the feature representation for traversing all clusters; K is the number of sample categories; τ2 is the cluster-level temperature coefficient; w a and are the feature representations of the a-th cluster in the image modality; v a and are the feature representations of the a-th cluster in the text modality; is the cluster-level multi-modal contrast loss from the image to the text direction; the total loss is the sum of the instance-level contrast loss, the cluster-level contrast loss, and the regularization loss.
[0025] In a second aspect, the present invention also provides a text-guided image clustering system, including:
[0026] A data acquisition module, configured to: acquire image data;
[0027] A clustering module, configured to: obtain an image clustering result according to the acquired image data and the image clustering model;
[0028] Among them, when training the image clustering model, text descriptions of the image data in the training set are generated, features of the images and texts are respectively extracted, text features are screened according to the similarity between each image clustering center and the text features, and the text features after screening are weighted to obtain the text features of each image; sample replacement is performed according to a preset similarity range, and the preset network model is trained according to the obtained image data and the replaced samples to obtain the image clustering model.
[0029] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the text-guided image clustering method described in the first aspect are implemented.
[0030] In a fourth aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor, and when the processor executes the program, the steps of the text-guided image clustering method described in the first aspect are implemented.
[0031] In a fifth aspect, the present invention also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the text-guided image clustering method described in the first aspect are implemented.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0033] 1. In the present invention, text features are screened according to the similarity between each image clustering center and the text features, the text features after screening are weighted to obtain the text features of each image, sample replacement is performed according to a preset similarity range, and the preset network model is trained according to the obtained image data and the replaced samples to obtain the image clustering model, which improves the accuracy of text and image matching and the clustering efficiency, and is applicable to the clustering of large-scale unlabeled image datasets;
[0034] 2. First, the present invention obtains an image dataset and generates corresponding text descriptions; then, performs word segmentation processing on the text descriptions and adds learnable prompts; then, extracts the feature information of the images and texts; next, inputs the text features into the text screening module to screen out some texts that are most similar to each image clustering center, and obtains the text features of each image through weighted calculation; subsequently, updates the image and text feature data through the random nearest neighbor sample selection module; then, inputs the updated image and text features into the instance-level mapping head for dimension adjustment and optimization; finally, trains through the cluster-level clustering head and minimizes the cluster-level contrast loss to achieve image clustering. By adopting the technical solution of the present invention, the accuracy and efficiency of image clustering can be effectively improved, and the accuracy and stability of the clustering results can be enhanced. Brief Description of the Drawings
[0035] The attached drawings forming a part of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions thereof of this embodiment are used to explain this embodiment and do not constitute an improper limitation to this embodiment.
[0036] Figure 1 It is a flowchart for implementing the method of Embodiment 1 of the present invention;
[0037] Figure 2 It is a general network framework diagram of Embodiment 1 of the present invention;
[0038] Figure 3 It is a framework diagram of the text screening module of Embodiment 1 of the present invention;
[0039] Figure 4 It is a framework diagram of the random nearest neighbor sample selection module of Embodiment 1 of the present invention;
[0040] Figure 5 It is a clustering effect diagram of Embodiment 1 of the present invention. Detailed Description of the Embodiment
[0041] The present invention will be further described below in conjunction with the drawings and embodiments.
[0042] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations for this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0043] Embodiment 1:
[0044] In today's data-driven era, unsupervised learning methods play an important role in the field of data mining and analysis. As a typical unsupervised learning means, clustering focuses on completing cluster division in the context of unlabeled data based on the similarity between data objects. Clustering methods are widely used in fields such as image segmentation and object detection, text data analysis, market segmentation, and bioinformatics.
[0045] However, traditional clustering methods have limitations in dealing with modern complex data. Taking the classic K-Means algorithm as an example, based on a simple distance measurement method, it is difficult to handle high-dimensional data and complex non-linear relationships, and it is difficult to deeply explore the potential structure of the data, resulting in insufficient accuracy and reliability of the clustering results.
[0046] In recent years, the rapid development of deep learning has made deep clustering a new research trend. Especially by using deep neural networks for automatic feature learning, image clustering algorithms can effectively extract more representative features from original images, thereby improving the clustering effect. However, traditional deep clustering methods still face some challenges, such as how to make full use of the knowledge of large-scale datasets, how to better capture the semantic relationships between images in an unsupervised environment, and how to improve the robustness and generalization ability of the model.
[0047] In this context, pre-trained models and contrastive learning methods have been widely applied in the field of image representation learning. Pre-trained contrastive learning models (such as CLIP, SimCLR, etc.) effectively learn rich semantic information in image data by maximizing the similarity between similar image pairs and minimizing the similarity between dissimilar image pairs. The application of these pre-trained models not only improves the quality of image representation but also significantly enhances the performance of image clustering. Therefore, combining pre-trained models and deep clustering methods can improve the quality of feature expression and enhance the model's ability to process high-dimensional complex data, providing stronger support for clustering tasks.
[0048] Based on this, this embodiment provides a text-guided image clustering method, which can improve the efficiency and accuracy of video retrieval; combines natural language processing technology, and further enhances the model's semantic understanding ability by combining the semantic information of images and text descriptions. Text guidance can help the model use external knowledge for supplementation and correction when the image content is unclear or has high ambiguity, improving the accuracy of clustering. In this embodiment, combining the pre-trained contrastive learning model and the text description-guided image clustering method can not only solve the problems of insufficient feature expression and insufficient semantic understanding in traditional methods but also provide a more efficient and accurate clustering solution for large-scale unlabeled image datasets, promoting the wide application of image clustering technology in multiple fields.
[0049] As Figure 1 shown, this embodiment may include:
[0050] Step S1, obtaining an image dataset; optionally, it can be implemented by a camera, a camera, and other electronic devices that can collect images.
[0051] Step S2, generating a text description for the image dataset obtained in step S1.
[0052] Step S3, performing word segmentation processing on the obtained text description and adding learnable prompts.
[0053] Step S4: Extract the feature information of the images and text. Optionally, the images and text can be respectively input into the image encoder and text encoder of the frozen CLIP (Contrastive Language-Image Pre-training) to obtain the image feature information and text feature information; or it can be achieved through the encoders of other models.
[0054] Step S5: Screen the text features of each image. Optionally, input the text feature information obtained in Step S4 into Figure 3 the text screening module shown in the figure, screen out some texts that are most similar to the clustering center of each image, and obtain the text features of each image through weighted calculation.
[0055] Step S6: Input the image features obtained in Step S4 and the text feature information obtained in Step S5 into the random neighbor sample selection module, delimit the range for each sample according to the similarity and replace the samples, and update the feature data.
[0056] Step S7: Input the image features obtained in Step S4 and the text features after being processed in Step S6 into the instance-level mapping head, adjust the dimensions and optimize the network and learnable prompts through the instance-level contrast loss;
[0057] Step S8: Input the image and text features in Step S4 and the image and text features obtained after Step S6 into the cluster-level clustering head, adjust the dimensions to match the number of clusters, and train the network with the cluster-level contrast loss.
[0058] Step S9: Input the test samples into the trained network, obtain the clustering results through the cluster-level contrast head, and complete the image clustering task.
[0059] In this embodiment, Step S2 includes the following contents:
[0060] Optionally, obtain the image dataset and generate the corresponding text descriptions. Input the image samples where N is the number of image samples, and X l represents the l-th image. In this embodiment, the pre-trained BLIP (Bootstrapped Language-Image Pretraining) model can be used to generate the necessary text descriptions for all images, and obtain where T l represents the text description generated from the images.
[0061] In this embodiment, Step S3 includes the following contents:
[0062] Optionally, the generated text descriptions Input it into CLIP.tokenize for tokenization, convert the text into the required tensor representation, and add learnable prompts to it. The optional specific form is Among them, is the token of the text description, p1, p2,... p s are the added learnable prompts.
[0063] In this embodiment, optionally, the step S4 includes the following steps:
[0064] Step S4.1: Input the image sample into the CLIP image encoder pre-trained based on ImageNet to obtain the image embedding feature
[0065] Step S4.2: Input the text after being tokenized in step S3 into the CLIP text encoder pre-trained based on ImageNet to obtain the preliminary text embedding feature
[0066] In this embodiment, the step S5 includes the following steps:
[0067] S5.1: Obtain the fine-grained image clustering center c j , j ∈ [1, k] ∩ Z.
[0068] S5.2: Select the top 5 text features with the highest similarity for each image clustering center through the formula where, j is the cluster referred to; sim is the cosine similarity; t is the text modality; l is the text embedding subscript;
[0069] is the l-th text embedding; c is the j-th clustering center; μ is the set selection threshold; arg top-μ is to calculate the similarity between the text embedding j and the clustering center c and select the top μ text embeddings with the highest similarity for each image clustering center, and store the subscripts in R j and store the subscripts in R j .
[0070] S5.3: In order to select the appropriate text for each image sample, calculate the similarity between each image sample and the M filtered texts, and use it as a weight factor, so as to calculate the final text embedding E by weighted calculation t .
[0071] In this embodiment, the step S6 includes the following steps:
[0072] S6.1. After the above operations, the images and texts of the entire batch will be processed by the Adapter to adjust the dimensions of the two-modal features.
[0073] S6.2. Calculate the similarity matrix S between samples in the same batch.
[0074] S6.3. For each sample Select the top 50 samples with the highest similarities from the l-th row of the similarity matrix S to form a candidate set Then, randomly select a sample from this candidate set and use to replace the feature with the l-th embedded feature in modality v, is the embedded feature of the selected sample, is the subscript of a sample randomly selected from the candidate set of the l-th sample. Among them, r represents the index of the randomly selected sample, and v represents the image or text modality. Through this process, we obtain the nearest neighbor feature matrix The whole process is as shown in the example of Figure 4 below.
[0075] In this embodiment, step S7 includes the following steps:
[0076] S7.1. Input the image feature E i obtained after the above processing and the text feature of the random nearest neighbor into two non-linear multi-layer perceptron (MLP) layers, the Adapter and the instance-level mapping head M ins to adjust the feature dimensions and obtain the two-modal feature matrices P and
[0077] S7.2. Use the feature matrices P and in step S71 for instance-level contrastive learning to minimize the instance-level loss L ins to achieve the alignment of the two-modal features. The instance-level contrastive loss is as follows:
[0078] f(z a , z b ) = sim(z a , z b );
[0079]
[0080] Among them, z a is the input image; z b is the input text embedding; N is the number of samples; b is the corresponding subscript when traversing all samples; τ1 is the instance-level temperature coefficient; and is the low-dimensional feature vector of the a-th sample pair; is the instance-level multi-modal contrast loss from the image to the text direction. The instance-level loss is the average of the losses in both the image-to-text and text-to-image directions.
[0081] In this embodiment, step S8 includes the following steps:
[0082] S8.1. Input the image and text features E i , E t and the image-text features of the random neighbors into two non-linear multi-layer perceptron (MLP) layers, Adapter and the cluster-level mapping head M clu to adjust the dimension to the number of clusters, and obtain the two-modal feature matrices W, and V,
[0083] S8.2. Use the feature matrices W, and V, from step 81 to perform cluster-level contrast learning, minimize the cluster-level loss L clu to adjust the network parameters and achieve clustering. The cluster-level contrast loss is as follows:
[0084]
[0085] where w b is the feature representation traversing all clusters; K is the number of sample categories; τ2 is the instance-level temperature coefficient; w a and are the feature representations of the a-th cluster in the image modality; v a and are the feature representations of the a-th cluster in the text modality; is the cluster-level multi-modal contrast loss from the image to the text direction. The cluster-level loss is also the average of the losses in both directions.
[0086] S8.3. Jointly optimize L ins and L clu , and at the same time add the regularization loss L nom to avoid the imbalance problem caused by most samples being assigned to a few categories. The regularization loss is as follows:
[0087]
[0088] where and are the means corresponding to the j-th cluster distribution in the image and text channels respectively.
[0089] In this embodiment, step S9 includes the following steps:
[0090] Input the test sample into the trained network, and obtain the probability distribution of each sample through the cluster-level contrastive head, so as to obtain the prediction result z = arg max M clu (E) to achieve the image clustering task; where M clu is the cluster-level contrastive head; E is the input image embedding feature.
[0091] In summary, in this embodiment, first, an image data set is obtained and corresponding text descriptions are generated; then, the text descriptions are tokenized and learnable prompts are added; next, the images and texts are respectively input into the frozen CLIP model to extract the feature information of the images and texts; then, the text features are input into the text screening module to screen out some texts that are most similar to each image clustering center, and the text features of each image are obtained through weighted calculation; subsequently, the image and text feature data are updated through the random neighbor sample selection module; then, the updated image and text features are input into the instance-level mapping head for dimension adjustment and optimization; finally, training is performed through the cluster-level clustering head, and the cluster-level contrastive loss is minimized to achieve image clustering. Through this embodiment, the accuracy and efficiency of image clustering can be effectively improved, and the accuracy and stability of the clustering results can be enhanced.
[0092] The following further illustrates the effect of the method in this embodiment through simulation:
[0093] Simulation experiment conditions:
[0094] The simulation experiment conditions of the present invention: Server GPU: NVIDIA GeForce RTX 3090, video memory 24GB.
[0095] The software platform of the simulation experiment of the present invention: Windows 11 Professional Edition, python3.9, torch + cu1212.3.1.
[0096] Simulation content and analysis of experimental results:
[0097] The simulation of this embodiment is to introduce text descriptions as a guide to perform image clustering work on the basis of the existing pre-trained contrastive learning large model CLIP. The data sets used in the simulation are: STL-10, CIFAR-10, CIFAR-100, ImageNet-10, ImageNet-Dogs.
[0098] The STL-10 dataset is a dataset designed specifically for small image classification tasks. It contains 10 categories, with 5,000 images in each category, for a total of 50,000 images. It was created by Fei et al. at Stanford University to evaluate the performance of small image classification models. Each image is 96×96 pixels in size and covers the following categories: airplane, bicycle, bird, cat, deer, dog, horse, ship, truck, and car.
[0099] The CIFAR dataset is a set of standard datasets widely used in computer vision tasks, mainly for the evaluation and development of image classification models. It includes three main versions: CIFAR-10 and CIFAR-100, each consisting of 32×32 color images.
[0100] (1) CIFAR-10: It contains 10 categories, with 6,000 pictures in each category, for a total of 60,000 images. These categories cover common objects such as animals and vehicles and are suitable for introductory research on tasks such as image classification and clustering.
[0101] (2) CIFAR-100: It has a structure similar to CIFAR-10 but is extended to 100 categories, covering 20 supercategories. Each category contains 600 pictures, offering higher diversity and challenge.
[0102] The ImageNet dataset is a widely used large-scale visual dataset, mainly applied to image classification and computer vision tasks. It contains over 14 million annotated images, covering objects in more than 21,000 categories. The annotation of the dataset is done manually, with hundreds to thousands of images in each category. These images are from the Internet and cover a variety of different objects and scenes, such as animals, plants, and buildings.
[0103] The simulation experiment compared the clustering results of 10 different clustering methods with the method of this embodiment (CLIP4Clustering) on the above 5 datasets. The comparison results are shown in Table 1:
[0104] Table 1 Comparison table of simulation experiment results on five datasets
[0105]
[0106] As can be seen from Table 1, compared with the baseline method, the method proposed in this embodiment achieved the best clustering effect on all datasets and all comparison methods, indicating that the pre-trained model and text description as a guide can effectively improve the clustering effect. To more intuitively display our clustering effect, in Figure 5The visualization results of the ImageNet-10 dataset at the 10th, 40th, and 80th rounds during the training process are shown. It can be found that as the training progresses, our method well realizes the division of samples of different categories.
[0107] Example 2:
[0108] This example provides a text-guided image clustering system, including:
[0109] A data acquisition module, configured to: obtain image data;
[0110] A clustering module, configured to: obtain an image clustering result according to the obtained image data and an image clustering model;
[0111] Among them, when training the image clustering model, generate text descriptions of the image data in the training set, extract the features of the images and texts respectively, screen the text features according to the similarity between each image clustering center and the text features, and obtain the text features of each image through a weighting method for the screened text features; perform sample replacement according to a preset similarity range, and train a preset network model according to the obtained image data and the replaced samples to obtain an image clustering model.
[0112] The working method of the system is the same as that of the text-guided image clustering method in Example 1, and will not be elaborated here.
[0113] Example 3:
[0114] This example provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the text-guided image clustering method described in Example 1 are implemented.
[0115] Example 4:
[0116] This example provides an electronic device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the program, the steps of the text-guided image clustering method described in Example 1 are implemented.
[0117] Example 5:
[0118] This example provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the text-guided image clustering method described in Example 1 are implemented.
[0119] The above are only the preferred embodiments of this embodiment and are not intended to limit this embodiment. For those skilled in the art, this embodiment can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this embodiment shall be included within the protection scope of this embodiment.
Claims
1. The image clustering method based on text guidance is characterized by: include: Get image data; According to the acquired image data and the image clustering model, an image clustering result is obtained; When the image clustering model is trained, a text description of the image data in the training set is generated, and the features of the image and text are extracted respectively. The text features are screened according to the similarity between each image cluster center and the text features, and the text features of each image are obtained by weighting the screened text features; samples are replaced according to the preset similarity range, and the preset network model is trained according to the acquired image data and the replaced samples to obtain the image clustering model.
2. The text-guided image clustering method according to claim 1, characterized in that: The text description is segmented, and the feature information of the image data and the text after the segmentation processing is extracted to obtain the image features and text feature information; the text feature information is screened to obtain the text features of each image; the image features and the screened text features are randomly selected as neighbor samples, and the range of each sample is defined according to the similarity; the extracted image features and the text features after sample selection are input into the preset mapping head, the dimension is adjusted and the network is optimized through the instance-level contrast loss; the extracted image features and text features, as well as the selected text features, are input into the clustering head, the dimension is adjusted to the cluster number matching state, and the network is trained with the cluster-level contrast loss.
3. The text-guided image clustering method according to claim 2, characterized in that: Calculate the similarity between each image sample and the preset number of texts after screening, use the similarity as a weight factor, and perform weighted calculation to obtain text embedding.
4. The text-guided image clustering method according to claim 3, characterized in that: The filtered text index set R j for: Among them, j is the cluster referred to; sim is the cosine similarity; t is the text modality; l is the text embedding subscript; is the lth text embedding; c j is the jth cluster center; μ is the selected threshold; arg top-μ is the text embedding and cluster center c j , and select the first μ text embeddings with the highest similarity as the image cluster center.
5. The text-guided image clustering method according to claim 2, characterized in that: Calculate the similarity matrix between samples in the same batch; for each sample, select a preset number of samples with the highest similarity from the rows of the similarity matrix to form a candidate set, and randomly select a sample from the candidate set for feature replacement.
6. The text-guided image clustering method according to claim 2, characterized in that: The instance-level contrast loss is: f(z a ,z b )=sim(z a ,z b ); Among them, z a is the incoming image; z b is the incoming text embedding; N is the number of samples; b is the corresponding subscript when traversing all samples; τ1 is the instance-level temperature coefficient; and is the low-dimensional feature vector of the a-th sample pair; It is an instance-level multimodal contrastive loss from image to text direction.
7. The text-guided image clustering method according to claim 6, characterized in that: The cluster-level contrast loss is: Joint Optimization L ins and L clu , while adding the regularization loss L nom , the regularization loss is: Among them, w b is the feature representation that traverses all clusters; K is the number of sample categories; τ2 is the cluster-level temperature coefficient; w a and is the feature representation of the ath cluster in the image modality; v a and is the feature representation of the a-th cluster in the text modality; is the cluster-level multimodal contrast loss from image to text direction; the total loss is the sum of instance-level contrast loss, cluster-level contrast loss, and regularization loss.
8. A text-guided image clustering system, characterized in that: include: The data acquisition module is configured to: acquire image data; The clustering module is configured to: obtain an image clustering result according to the acquired image data and the image clustering model; When the image clustering model is trained, a text description of the image data in the training set is generated, and the features of the image and text are extracted respectively. The text features are screened according to the similarity between each image cluster center and the text features, and the text features of each image are obtained by weighting the screened text features; samples are replaced according to the preset similarity range, and the preset network model is trained according to the acquired image data and the replaced samples to obtain the image clustering model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the program, the steps of the text-guided image clustering method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the text-guided image clustering method according to any one of claims 1 to 7 are implemented.