Automatic Image Category Annotation Method and Device Based on Density Peak Clustering Algorithm

By combining the convolutional autoencoder model and density peak clustering algorithm in the automatic image category annotation, the problems of unreasonable density measurement and low clustering efficiency in the prior art are solved, and efficient automatic category annotation of labelless image data is achieved.

CN115147632BActive Publication Date: 2025-05-27HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210800775.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2025-05-27
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

The existing density peak clustering methods have problems such as unreasonable density measurement, manual clustering center selection, and parameter sensitivity in image data processing, which are difficult to effectively apply to large-scale image data sets. In addition, deep learning methods carry out feature learning and clustering tasks in step, resulting in low clustering efficiency and unoptimized feature representation.

Method used

The image category automatic labeling method based on the density peak clustering algorithm is adopted, and the labelless image data is reduced and feature extracted through the convolutional autoencoder model. The candidate clustering center is selected in the feature vector space with the density peak clustering method, and the confidence vector is updated using the semi-supervised clustering method until the KL divergence value of the clustering result is less than the given threshold.

Benefits of technology

Automatic category labeling of label-free image data is realized, and the problems of manual labeling are solved, which are time-consuming, cost-effective, low accuracy and poor efficiency, and improve the clustering efficiency and optimization of feature representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147632B_ABST
    Figure CN115147632B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic image category annotation method and device based on a density peak clustering algorithm, including convolutional autoencoder model training and convolutional encoder-clustering joint training. The image dataset to be annotated is input into the model to train the convolutional autoencoder module, and then the trained convolutional encoder module is taken out to reduce the image data to a low-dimensional feature vector space. The low-dimensional feature vectors are input into the convolutional encoder-clustering joint training module, and the density peak clustering method is used in the feature vector space to select candidate clustering centers and find a high-confidence data set. The categories of the high-confidence data set are used as the true labels to train the convolutional encoder module to obtain a clustering result with high credibility. Finally, the input unlabeled image data is annotated with the feature vector categories. The present invention can achieve automatic category annotation for unlabeled image data, solving the problems of long time consumption, high cost, low accuracy, and poor efficiency in current manual category annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of clustering analysis, and particularly relates to an image category automatic annotation method and device based on a density peak clustering algorithm. Background Art

[0002] In traditional density-based clustering methods, the density peak clustering method is a simple and efficient algorithm, which is easy to execute and has high scalability, and is widely applied to various tasks. The density-based clustering method calculates the local density of data points through an appropriate density function, and based on this, finds the association between data points and clusters the data points. In the field of image clustering, traditional clustering methods usually first reduce the dimensionality of image data into low-level image feature representations for encoding, and then cluster the feature representations. Recent deep clustering methods apply deep learning to the clustering field, combine feature learning and clustering into the same model, use autoencoders or other networks to learn the information representation of images, and then use traditional clustering methods to cluster the learned information representation.

[0003] For the existing density peak clustering methods, there are problems such as unreasonable density measurement, manual selection of clustering centers, and parameter sensitivity. Since the metric method based on the norm distance cannot calculate the similarity between image data, it will amplify the defects of the algorithm when processing image data. Coupled with the limitations of computing resources such as CPU memory and computing time, it is difficult for traditional clustering methods to be directly applied to large-scale image data sets. Some traditional clustering methods will use dimensionality reduction methods to obtain low-level features of images for clustering, but such low-level feature representations are easily affected by changes in the image scene and object appearance, and image transformations such as rotation and brightness changes will also have a greater impact on feature learning. Many research methods use deep unsupervised feature learning methods to learn the feature representations of images, but most deep learning methods perform feature learning and clustering tasks step by step. Although the learned features can reconstruct the input data, they cannot be directly applied to the clustering algorithm to obtain clustering results. In addition, directly using the K-means algorithm to cluster the features will result in a slow convergence speed, high time complexity in the clustering part, and only guarantee local optimality. Moreover, the method of separating feature extraction and clustering makes it difficult for the network model to learn the optimal feature representation, affecting the final clustering effect. Summary of the Invention

[0004] In view of the above problems, the present invention provides an image category automatic annotation method and device based on a density peak clustering algorithm, which is used to automatically annotate the categories of unlabeled image data, and solves the problems of long time consumption, high cost, low accuracy, and poor efficiency of the current method for manually annotating pictures.

[0005] In a first aspect of the present invention, an automatic image category annotation method based on density peak clustering algorithm is provided, and the method includes:

[0006] Obtain an unlabeled image dataset of the category to be annotated;

[0007] Input the unlabeled image dataset into a convolutional autoencoder model for training. The specific process includes: using the convolutional encoder module in the convolutional autoencoder model to reduce and compress the input unlabeled image data into a low-dimensional feature vector, and then using the convolutional decoder module to recover the image from the low-dimensional feature vector to obtain reconstructed image data. Calculate the reconstruction loss between the input unlabeled image data and the reconstructed image data. When the reconstruction loss is less than a given threshold, it is determined that the training of the convolutional autoencoder model ends;

[0008] Retain the convolutional encoder module in the trained convolutional autoencoder model, and use the convolutional encoder module to obtain a set of low-dimensional feature vectors of the unlabeled image dataset;

[0009] Input the set of low-dimensional feature vectors into a convolutional encoder-clustering joint training module for joint training. The specific process includes: using the density peak clustering method to calculate the local density of the feature vector points in the set of low-dimensional feature vectors and the distance from the high-density points, and multiplying the local density of the feature vector points and the distance from the high-density points to obtain the gamma value of the feature vector; sort the gamma values of all feature vectors in the set of low-dimensional feature vectors in descending order, select the first m feature vector points as candidate cluster centers to obtain a set of candidate cluster centers; calculate the distances from the remaining feature vector points to each candidate cluster center to obtain an m-dimensional distance vector; take the reciprocal of each component in the m-dimensional distance vector and normalize it to obtain an m-dimensional category confidence vector, and take the column where the component with the largest value in the m-dimensional category confidence vector is located as the true category label of the feature vector point to obtain the clustering result of the category confidence vector. Use the true category label to train the convolutional encoder module with labels; update the confidence vector matrix of the remaining feature vector points with the trained convolutional encoder module until the KL divergence value between the clustering results of two consecutive rounds is less than a given threshold, and the training ends;

[0010] Take the clustering result obtained after the training ends as the final clustering result, and use the final clustering result to annotate the input unlabeled images to obtain the final annotated image dataset.

[0011] A further technical solution of the present invention is: before inputting the unlabeled image dataset into the convolutional autoencoder model for training, first perform data augmentation on the input unlabeled image data and add random Gaussian noise.

[0012] A further technical solution of the present invention is: calculate the reconstruction loss between the input unlabeled image data and the reconstructed image data, and the specific expression is:

[0013]

[0014] Among them, n represents the size of the input unlabeled image dataset, and X i represents the input unlabeled image data sample, represents X i is the reconstructed image data obtained through the convolutional encoder module and the convolutional decoder module. φ represents the parameters of the convolutional encoder module, θ represents the parameters of the convolutional decoder module, and f φ represents the mapping from the input unlabeled image data to the feature vector implemented by the convolutional encoder module, and g θ represents the mapping from the feature vector to the reconstructed image data implemented by the convolutional decoder module, and L rec represents the reconstruction loss of the entire convolutional autoencoder model.

[0015] A further technical solution of the present invention is: using the density peak clustering method to calculate the local density and the distance from the high-density points of the feature vector points in the low-dimensional feature vector set. The specific method includes:

[0016] Calculate the distance from each feature vector point to its k nearest neighbors, and calculate the mean μ and standard deviation σ of the k-nearest neighbor distances;

[0017]

[0018]

[0019] where d(x, x i ) represents the Euclidean distance from the feature vector point x to its k nearest neighbor x i ;

[0020] According to the Raida criterion, an upper limit θ = μ + 3σ is calculated. Traverse the k-nearest neighbor distances, remove the neighbors greater than the upper limit θ, and obtain a new k-nearest neighbor set. Calculate the local density ρ of the data points according to the updated k-nearest neighbor set as:

[0021]

[0022] where the updated k-nearest neighbor set is AKNN = {x j |d(x, x j ) ≤ d(x, x k ) ∧ d(x, x j ) ≤ θ = μ + 3σ}, x represents the feature vector point, x j is an object in the k nearest neighbors of x, x j ∈AKNN means that x j belongs to the k nearest neighbors of x, and d(x, x j ) represents the distance between x and x jThe Euclidean distance, where the symbol ^ represents conditional AND;

[0023] The distance δ from the eigenvector dot to the high-density point is:

[0024]

[0025] where ρ i represents the local density of point i, D represents the set of all eigenvector points, and d(x i , x j ) represents the Euclidean distance between x i and x j two points.

[0026] In the second aspect of the present invention, an automatic image category annotation device based on the density peak clustering algorithm is provided. The device includes:

[0027] An image acquisition unit for acquiring an unlabeled image data set of the category to be annotated;

[0028] A convolutional autoencoder model training unit for inputting the unlabeled image data set into the convolutional autoencoder model for training. The specific process includes: using the convolutional encoder module in the convolutional autoencoder model to reduce the dimensionality of the input unlabeled image data to a low-dimensional eigenvector, and then using the convolutional decoder module to recover the low-dimensional eigenvector into reconstructed image data, calculating the reconstruction loss between the input unlabeled image data and the reconstructed image data, and determining that the training of the convolutional autoencoder model stops when the reconstruction loss is less than a given threshold;

[0029] A low-dimensional eigenvector set acquisition unit for retaining the convolutional encoder module in the trained convolutional autoencoder model and using the convolutional encoder module to acquire the low-dimensional eigenvector set of the unlabeled image data set;

[0030] A convolutional encoder-clustering joint training module training unit for inputting the low-dimensional eigenvector set into the convolutional encoder-clustering joint training module for joint training. The specific process includes: using the density peak clustering method to calculate the local density and the distance from the high-density point of the eigenvector points in the low-dimensional eigenvector set, and multiplying the local density and the distance from the high-density point of the eigenvector points to obtain the gamma value of the eigenvector; sorting the gamma values of all eigenvectors in the low-dimensional eigenvector set in descending order, selecting the first m eigenvector points as candidate clustering centers to obtain a candidate clustering center set; calculating the distances from the remaining eigenvector points to each candidate clustering center to obtain an m-dimensional distance vector; and obtaining an m-dimensional category confidence vector by taking the reciprocal of each component in the m-dimensional distance vector and normalizing it.

[0031] Take the column where the component with the largest value in the m-dimensional class assignment confidence vector as the true class label of the feature vector point, obtain the clustering result of the class assignment confidence vector, and use the true class label as the labeled training convolutional encoder module; update the confidence vector matrix of the remaining feature vector points with the trained convolutional encoder module until the KL divergence value between the clustering results of two consecutive rounds is less than the given threshold, and the training ends;

[0032] An image dataset annotation unit, which is used to take the clustering result obtained after the training ends as the final clustering result, and label the input unlabeled images with the class division result to obtain the final annotated image dataset.

[0033] A further technical solution of the present invention is that before inputting the unlabeled image dataset into the convolutional autoencoder model for training, the convolutional autoencoder model training unit first performs data augmentation on the input unlabeled image data and adds random Gaussian noise.

[0034] In a third aspect of the present invention, there is provided an image class automatic annotation device based on a density peak clustering algorithm, including: a processor; and a memory, wherein a computer executable program is stored in the memory, and when the computer executable program is executed by the processor, the above-mentioned image class automatic annotation method based on the density peak clustering algorithm is executed.

[0035] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the processor is made to execute the above-mentioned image class automatic annotation method based on the density peak clustering algorithm.

[0036] The present invention proposes an image class automatic annotation method, device and storage medium based on a density peak clustering algorithm. The method mainly includes a convolutional autoencoder model pre-training module and a convolutional encoder-clustering joint training module. The image dataset to be annotated is input into the pre-training module to train the convolutional encoder module, and then the trained convolutional encoder module is taken out to reduce the image data to a low-dimensional feature vector space; the low-dimensional feature vectors are input into the convolutional encoder-clustering joint training module, and the density peak clustering method is used in the feature vector space to select candidate clustering centers and find a high-confidence data set. Using the semi-supervised clustering method, the class of the high-confidence data set is used as the true label to train the convolutional encoder, and finally a clustering result with high credibility is obtained. Finally, the input unlabeled image data is class-annotated with the feature vector class. The method of the present invention can realize automatic class annotation of unlabeled image data, and solves the problems of long time-consuming, high cost, low accuracy and poor efficiency of the current method of manually annotating pictures. Description of the Drawings

[0037] Figure 1It is a schematic flowchart of the method for automatically annotating image categories based on the density peak clustering algorithm in the embodiments of the present invention;

[0038] Figure 2 It is a schematic diagram of the method for training a convolutional autoencoder model in the embodiments of the present invention;

[0039] Figure 3 It is a schematic diagram of the convolutional encoder-clustering joint training method in the embodiments of the present invention;

[0040] Figure 4 It is a schematic structural diagram of the device for automatically annotating image categories based on the density peak clustering algorithm in the embodiments of the present invention;

[0041] Figure 5 It is the architecture of the computer device in the embodiments of the present invention;

[0042] Figure 6 It is the distribution diagram of feature vectors in the embodiments of the present invention;

[0043] Figure 7 It is a schematic diagram of partial clustering results of the MNIST dataset in the embodiments of the present invention. Detailed implementation manners

[0044] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0045] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0046] In the description of the present invention, "a plurality" and "several" mean at least two, such as two, three, etc., unless otherwise specifically defined.

[0047] The embodiments of the present invention provide the following embodiments for an image category automatic annotation method, device, and storage medium based on the density peak clustering algorithm:

[0048] Based on Embodiment 1 of the present invention

[0049] This embodiment is used to illustrate the training part of the convolutional autoencoder model in the image category automatic annotation method based on the density peak clustering algorithm. As Figure 1 shown, it is the flowchart of the image category automatic annotation method based on the density peak clustering algorithm according to the embodiment of the present invention:

[0050] Obtain an unlabeled image dataset of the category to be annotated;

[0051] Input the unlabeled image dataset into the convolutional autoencoder model for training. As Figure 2 shown, it is the training process of the convolutional autoencoder model. The specific process includes: using the convolutional encoder module in the convolutional autoencoder model to reduce the dimensionality of the input unlabeled image data and compress it into a low-dimensional feature vector, and then using the convolutional decoder module to restore the low-dimensional feature vector to obtain the reconstructed image data. Calculate the reconstruction loss between the input unlabeled image data and the reconstructed image data. When the reconstruction loss is less than a given threshold, it is determined that the training of the convolutional autoencoder model is terminated;

[0052] In the specific implementation process, the convolutional neural network uses convolutional kernels to extract the features of image data. By stacking multiple convolutional layers, the ability of the convolutional neural network model to extract features and the amount of information obtained can be increased. Using multiple output channels can extract different levels of image information respectively and enhance the representation ability of features. However, if convolutional layers are simply stacked to extract features, it cannot be guaranteed that the spatial distribution of the feature vectors output by the network is consistent with the true category distribution of the original image data, which will lead to inaccurate clustering results of the subsequent clustering algorithm on the feature vectors. The general training method of a convolutional neural network model is to input data into the network model to obtain the model classification result, use a loss function to calculate the error between the known category label and the model output category, and use the error to update the network weights. In the application scenario of this embodiment, the input dataset is unlabeled, and the true category label cannot be used to calculate the loss function to update the network weights. Therefore, the embodiment uses an autoencoder structure to train the convolutional autoencoder model, using the "encoder-decoder" network structure, which avoids the need for sample labels in traditional network models. The convolutional autoencoder model uses the convolutional encoder module to reduce the dimensionality and extract features of the input image data to obtain a one-dimensional feature vector, and then uses the convolutional decoder to perform picture restoration operations on the one-dimensional feature vector to obtain the reconstructed picture data. The convolutional autoencoder model evaluates the training effect of the model by calculating the difference between the reconstructed picture and the input picture. If the gap between the reconstructed picture and the input picture is too large, it means that the training effect of the convolutional autoencoder model is not good enough; if the gap between the reconstructed picture and the input picture is very small, it means that the model training effect is good.

[0053] Furthermore, calculate the reconstruction loss between the input unlabeled image data and the reconstructed image data. The specific expression is:

[0054]

[0055] Among them, n represents the size of the input unlabeled image dataset, and X i represents the input unlabeled image data sample, represents the reconstructed image data obtained by X i through the convolutional encoder module and the convolutional decoder module. φ represents the parameters of the convolutional encoder module, θ represents the parameters of the convolutional decoder module, and f φ represents the mapping from the input unlabeled image data to the feature vector implemented by the convolutional encoder module, and g θ represents the mapping from the feature vector to the reconstructed image data implemented by the convolutional decoder module, and L rec represents the reconstruction loss of the entire convolutional autoencoder model. The smaller the value of the reconstruction loss L rec , the better the training effect of the model on the entire input image dataset, and the better the representation effect of the feature vector extracted by its encoder.

[0056] Furthermore, before inputting the unlabeled image dataset into the convolutional autoencoder model for training, the input unlabeled image data is first subjected to data augmentation and random Gaussian noise is added. Correspondingly, data augmentation operations are performed on the input image, operations such as (translation, rotation, flipping) are performed on the image, and random noise is applied to improve the robustness of model training. The calculation method of the loss function for enhancing the convolutional autoencoder training model is:

[0057]

[0058] In the specific implementation process, simple data augmentation operations are performed on the input image data, such as image translation, image rotation, horizontal / vertical flipping, etc., and random Gaussian noise is added. The convolutional autoencoder model is trained with the image data after data augmentation. The convolutional encoder part reduces the image data after data augmentation to the low-dimensional feature vector space, and then the convolutional decoder restores the image through the feature vector to obtain the reconstructed image data. The reconstruction loss between the input image data and the reconstructed image data is calculated. If the reconstruction loss is greater than the given threshold or the training cutoff condition is not met, the convolutional autoencoder model continues to be trained with the input image data. If the reconstruction loss is less than the given threshold, the model training is stopped, and the trained convolutional autoencoder model is saved.

[0059] Based on Embodiment 2 of the present invention

[0060] This embodiment is based on Embodiment 1 and is used to illustrate the convolutional encoder-clustering joint training module part in the image category automatic annotation method based on the density peak clustering algorithm on the basis of Embodiment 1.

[0061] Retain the convolutional encoder module in the trained convolutional autoencoder model, and use the convolutional encoder module to obtain a set of low-dimensional feature vectors of the unlabeled image dataset;

[0062] In the specific implementation process, Example 1 uses an unlabeled picture dataset to be annotated to train the convolutional autoencoder model, which can enable the convolutional encoder to extract the features of the original picture dataset well and obtain a set of low-dimensional feature vectors. After obtaining the set of low-dimensional feature vectors, the density peak clustering method is used to cluster the feature vector set, and the feature vector set is classified.

[0063] In the specific implementation process, discard the convolutional decoder module, retain the convolutional encoder module in the trained convolutional autoencoder model, and reduce the dimension of the image dataset to be annotated into the low-dimensional feature vector space. Input the set of low-dimensional feature vectors into the convolutional encoder-clustering joint training module.

[0064] Input the set of low-dimensional feature vectors into the convolutional encoder-clustering joint training module for joint training. As Figure 3 shown, it is the convolutional encoder-clustering joint training process. The specific process includes: using the density peak clustering method to calculate the local density and the distance from the high-density point of the feature vector points in the set of low-dimensional feature vectors, and multiplying the local density and the distance from the high-density point of the feature vector points to obtain the gamma value of the feature vector; sorting the gamma values of all feature vectors in the set of low-dimensional feature vectors in descending order, and selecting the first m feature vector points as candidate clustering centers to obtain a set of candidate clustering centers; calculating the distances from the remaining feature vector points to each candidate clustering center to obtain an m-dimensional distance vector; taking the reciprocal of each component in the m-dimensional distance vector and normalizing it to obtain an m-dimensional class confidence vector, and taking the column where the component with the largest value in the m-dimensional class confidence vector is located as the true class label of the feature vector point to obtain the clustering result of the class confidence vector, and using the true class label to train the convolutional encoder module with labels; updating the confidence vector matrix of the remaining feature vector points with the trained convolutional encoder module until the KL divergence value between the clustering results of the previous and next rounds is less than a given threshold, and the training ends;

[0065] In the preferred embodiment, first, the local density of the feature vector and the distance from the high-density point are calculated according to the density peak method, and the two are multiplied to obtain the gamma value of the feature vector. The gamma values are sorted in descending order, and the feature vector points at the front positions are selected as candidate clustering centers to obtain a set of candidate clustering centers. Calculate the distances from the remaining feature vector points to each candidate clustering center, and obtain an overall confidence vector set. The category to which the feature vector points with high confidence belong is used as their true category label, and the convolutional encoder module is trained with the labeled training data. The confidence vectors of the remaining points are updated with the trained convolutional encoder module. By calculating the KL divergence value between the clustering results of two consecutive rounds, if it is greater than the given threshold, repeat the above training method, first train the network with the high-confidence feature vector set, and then update the remaining data points with the trained network; if it is less than the given threshold, the training ends.

[0066] The clustering result obtained after the training ends is used as the final clustering result, and the input unlabeled images are labeled with the final clustering result to obtain the final labeled image dataset.

[0067] The steps of the traditional density peak clustering method are as follows: Use adaptive k-nearest neighbors to calculate the local density of each data point. The commonly used local density calculation method estimates and calculates using the k-nearest neighbor information of the data point, and the commonly used calculation formula is The greater the local density ρ of the data point, the higher the local density of the point, and the greater the possibility that it belongs to the true clustering center. If the k nearest neighbors of the data point are directly calculated without screening, the k-nearest neighbor set is likely to contain boundary points, outliers, or points of other classes. At this time, the calculated local density may have a negative impact on the class assignment of the current data point, resulting in continuous errors in subsequent object division. To solve this problem, the present invention proposes the concept of adaptive k-nearest neighbors and calculates the local density of the data point accordingly.

[0068] Specifically, the density peak clustering method is used to calculate the local density of the feature vector points in the low-dimensional feature vector set and the distance from the high-density point. The specific method includes:

[0069] Calculate the distances from each feature vector point to its k nearest neighbors, and calculate the mean μ and standard deviation σ of the k-nearest neighbor distances;

[0070]

[0071]

[0072] where d(x,x i ) represents the Euclidean distance from the feature vector point x to its k nearest neighbor x i ;

[0073] According to the 3-sigma rule, an upper limit θ = μ + 3σ is calculated. Traverse the k-nearest neighbor distances, remove the neighbors whose distances are greater than the upper limit θ, and obtain a new set of k-nearest neighbors. Calculate the local density ρ of the data point based on the updated set of k-nearest neighbors as follows:

[0074]

[0075] where the updated set of k-nearest neighbors is AKNN = {x j | d(x, x j ) ≤ d(x, x k ) ∧ d(x, x j ) ≤ θ = μ + 3σ}, x represents the feature vector point, x j is an object among the k-nearest neighbors of x, x j ∈ AKNN means that x j belongs to the k-nearest neighbors of x, d(x, x j ) represents the Euclidean distance between x and x j , and the symbol ^ represents conditional AND;

[0076] The distance δ from the feature vector point to the high-density point is:

[0077]

[0078] where ρ i represents the local density of point i, D represents the set of all feature vector points, and d(x i , x j ) represents the Euclidean distance between x i and x j .

[0079] In the specific implementation process, the density peak clustering method believes that the characteristics of the clustering center are: its own local density is high, and the distance from other high-density points is far. Therefore, the gamma value obtained by multiplying the local density ρ by the distance δ is used as the judgment criterion for whether a data sample is a clustering center. The larger the gamma value of the data sample, the more likely it is to be a clustering center. According to this criterion, the calculated gamma values are sorted in descending order, and the data samples in the front positions are preferentially selected as clustering centers. Select the first m data points as the clustering centers of the input data set, and represent them with m-dimensional vectors e 1 , e 2 ,..., e m , where e i means that the i-th position of the vector is 1 and the other positions are 0. Traverse the remaining data sample points, calculate the distance from each data point to the clustering center, and obtain an m-dimensional distance vector [d 1 , d 2 ,..., d m , and take the reciprocal of each component in the distance vector and normalize it to obtain an m-dimensional vector [p1 , p 2 ,..., p m as the class allocation confidence, i.e., p i represents the confidence that the current data point belongs to the i-th class.

[0080] Further, to prevent the randomness and errors of a single clustering result, the embodiments of the present invention use the sample classes with high confidence in each round as pseudo-labels and train the convolutional encoder module using a semi-supervised method. First, the input image data is reduced to the feature vector space through the convolutional encoder module. In the feature vector space, the density peak clustering method is applied. The local density of each data point is calculated using adaptive k-nearest neighbors. The closest distance from the data point to the high-density points is calculated through the local density. The gamma values obtained by multiplying the local density and the closest distance are sorted in descending order, and the top m data points are taken as candidate clustering centers. The distances from the remaining data points to each candidate clustering center are calculated to obtain the confidence vector set of the entire data set.

[0081] The data points with high confidence are taken out, and the class to which they belong is used as the true class label to obtain the sure class data set. This part of the data is used to retrain the convolutional encoder module to update the convolutional encoder module. Then, the updated convolutional encoder module is used to calculate and update the confidence vectors of the remaining data points, and the column where the component with the largest value in the confidence vector is located is used as the class label of the data point. Based on the above process, the purpose of training the convolutional encoder module with data points with high confidence and updating the weights is achieved, making the class labels more reliable. When the difference between the last two clustering results obtained by the convolutional encoder module is less than a given threshold, the training ends, and the clustering result at this time is used as the final class of the data points, that is, the final class of the input unlabeled image data set.

[0082] Based on Embodiment 3 of the present invention

[0083] Hereinafter, with reference to Figure 4Describe the apparatus corresponding to the method according to Embodiments 1-3 of the present disclosure. An automatic image category annotation apparatus 400 based on a density peak clustering algorithm includes an image acquisition unit 401 for acquiring an unlabeled image data set of the category to be annotated; a convolutional autoencoder model training unit 402 for inputting the unlabeled image data set into the convolutional autoencoder model for training. The specific process includes: using the convolutional encoder module in the convolutional autoencoder model to reduce the dimensionality of the input unlabeled image data to a low-dimensional feature vector, and then using the convolutional decoder module to recover the image from the low-dimensional feature vector to obtain reconstructed image data, calculating the reconstruction loss between the input unlabeled image data and the reconstructed image data, and determining the end of the convolutional autoencoder model training when the reconstruction loss is less than a given threshold; a low-dimensional feature vector set acquisition unit 403 for retaining the convolutional encoder module in the trained convolutional autoencoder model and using the convolutional encoder module to obtain a low-dimensional feature vector set of the unlabeled image data set; a convolutional encoder-clustering joint training module training unit 404 for inputting the low-dimensional feature vector set into the convolutional encoder-clustering joint training module for joint training. The specific process includes: using the density peak clustering method to calculate the local density and the distance to the high-density point of the feature vector points in the low-dimensional feature vector set, and multiplying the local density and the distance to the high-density point of the feature vector points to obtain the gamma value of the feature vector; sorting the gamma values of all feature vectors in the low-dimensional feature vector set in descending order, selecting the first m feature vector points as candidate cluster centers to obtain a candidate cluster center set; calculating the distances from the remaining feature vector points to each candidate cluster center to obtain an m-dimensional distance vector; taking the reciprocal of each component in the m-dimensional distance vector and normalizing it to obtain an m-dimensional category assignment confidence vector, and taking the column where the component with the largest value in the m-dimensional category assignment confidence vector is located as the true category label of the feature vector point to obtain the clustering result of the category assignment confidence vector, and using the true category label to train the convolutional encoder module with labels; updating the confidence vector matrix of the remaining feature vector points with the trained convolutional encoder module until the KL divergence value between the clustering results of the previous and next rounds is less than a given threshold, and the training ends; an annotated image data set unit 405 for using the clustering result obtained after the training ends as the final clustering result and annotating the input unlabeled image with the category division result to obtain the final annotated image data set. In addition to the above five units, the apparatus 400 may further include other components. However, since these components are not related to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.

[0084] Further, before inputting the unlabeled image data set into the convolutional autoencoder model for training, the convolutional autoencoder model training unit 402 first performs data augmentation on the input unlabeled image data and adds random Gaussian noise.

[0085] The specific working process of the image category automatic annotation device 400 based on the density peak clustering algorithm may refer to the descriptions of Embodiments 1-3 of the image category automatic annotation method based on the density peak clustering algorithm above, and will not be elaborated here.

[0086] Based on Embodiment 4 of the present invention

[0087] The device according to the embodiment of the present invention can also be implemented with the aid of Figure 5 the architecture of the computing device shown. Figure 5 The architecture of the computing device is shown. As Figure 5 shown, a computer system 501, a system bus 503, one or more CPUs 504, an input / output 502, a memory 505, etc. The memory 505 can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the methods of Embodiments 1-3. Figure 5 The architecture shown is only exemplary. When implementing different devices, adjust one or more components in Figure 5 according to actual needs.

[0088] Based on Embodiment 5 of the present invention

[0089] The embodiment of the present invention can also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium according to Embodiment 5. When the computer-readable instructions are run by a processor, the above-mentioned image category automatic annotation method based on the density peak clustering algorithm according to Embodiments 1-3 of the present invention can be executed with reference to the above drawings.

[0090] For the above-mentioned image category automatic annotation method, device and storage medium based on the density peak clustering algorithm, the embodiment of the present invention selects the MNIST dataset and the USPS dataset to test the performance of the method of the present invention, and selects K-means, DPC, DEC, DCN as comparison methods.

[0091] The evaluation index selects the clustering accuracy (ACC), and the calculation formula is

[0092]

[0093] where n represents the number of samples in the dataset, y represents the true label of the sample, and y' represents the clustering label.

[0094]

[0095] First, the convolutional autoencoder network is pre-trained using the MNIST dataset and the USPS dataset respectively, and then the proposed method for class labeling of unlabeled image datasets is tested on the MNIST-TSET dataset and the USPS dataset. Table 1 shows the comparison of clustering accuracies between the method of the present invention and other methods on the MNIST-TEST and USPS datasets:

[0096] Table 1 Comparison of accuracies between the present method and other methods on different datasets

[0097]

[0098] As can be seen from the table, the clustering performance of deep clustering methods (DEC, DCN) is significantly better than that of traditional clustering methods (K-means, DPC). The method of the present invention (OUR) trains the convolutional autoencoder using the image data after data augmentation, which improves the robustness of the network. The convolutional encoder is used to reduce the dimensionality of the image data, enabling traditional clustering methods to handle the feature vectors well and making full use of the advantage of high scalability of traditional clustering methods. Subsequently, the dataset is divided into a high-confidence data set and a low-confidence data set, and the high-confidence data set is used as the true label. A semi-supervised training method is adopted to train the network, update the weights and the confidence vector, which can enable the convolutional encoder to better learn the clustering feature vectors and improve the final clustering accuracy. Figure 6 Visualizing the feature vectors after dimensionality reduction by the convolutional encoder, it can be seen that through semi-supervised joint training, the method of the present invention can well group the data of the same class together and separate it from the data of different classes. Figure 7 Shows partial clustering results of the MNIST dataset.

[0099] Using Examples 1-5 and the above performance analysis, the method of the present invention can achieve automatic class labeling of unlabeled image data, solving the problems of long time consumption, high cost, low accuracy, and poor efficiency in the current method of manually labeling pictures. The method mainly includes a pre-training module of the convolutional autoencoder model and a joint training module of the convolutional encoder-clustering. The unlabeled image dataset to be labeled is input into the pre-training module to train the convolutional autoencoder module, and then the trained convolutional encoder module is taken out to reduce the dimensionality of the image data to a low-dimensional feature vector space; the low-dimensional feature vectors are input into the joint training module of the convolutional encoder-clustering. The density peak clustering method is used in the feature vector space to select candidate clustering centers and find the high-confidence data set. The semi-supervised clustering method is used to train the convolutional encoder with the class of the high-confidence data set as the true label, and finally a clustering result with high credibility is obtained. Finally, the input unlabeled image data is class-labeled with the class of the feature vectors.

[0100] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An automatic image category annotation method based on density peak clustering algorithm, characterized in that, the method includes: Obtain an unlabeled image dataset of the category to be annotated; Input the unlabeled image dataset into the convolutional autoencoder model for training. The specific process includes: using the convolutional encoder module in the convolutional autoencoder model to reduce and compress the input unlabeled image data into a low-dimensional feature vector, and then using the convolutional decoder module to recover the low-dimensional feature vector into reconstructed image data, calculating the reconstruction loss between the input unlabeled image data and the reconstructed image data. When the reconstruction loss is less than a given threshold, it is determined that the training of the convolutional autoencoder model ends; Retain the convolutional encoder module in the trained convolutional autoencoder model, and use the convolutional encoder module to obtain a set of low-dimensional feature vectors of the unlabeled image dataset; Input the set of low-dimensional feature vectors into the convolutional encoder-clustering joint training module for joint training. The specific process includes: using the density peak clustering method to calculate the local density and the distance to the high-density point of the feature vector points in the set of low-dimensional feature vectors, and multiplying the local density and the distance to the high-density point of the feature vector points to obtain the gamma value of the feature vector; sort the gamma values of all feature vectors in the set of low-dimensional feature vectors in descending order, select the first m feature vector points as candidate cluster centers to obtain a set of candidate cluster centers; calculate the distances from the remaining feature vector points to each candidate cluster center to obtain an m-dimensional distance vector; take the reciprocal of each component in the m-dimensional distance vector and normalize it to obtain an m-dimensional category confidence vector, and take the column where the component with the largest value in the m-dimensional category confidence vector is located as the true category label of the feature vector point to obtain the clustering result of the category confidence vector, and use the true category label to train the convolutional encoder module with labels; update the confidence vector matrix of the remaining feature vector points with the trained convolutional encoder module until the KL divergence value between the clustering results of the previous and next rounds is less than a given threshold, and the training ends; Take the clustering result obtained after the training ends as the final clustering result, and use the final clustering result to annotate the input unlabeled images to obtain the final annotated image dataset.

2. The automatic image category annotation method based on density peak clustering algorithm according to claim 1, characterized in that, Before inputting the unlabeled image dataset into the convolutional autoencoder model for training, first perform data augmentation on the input unlabeled image data and add random Gaussian noise.

3. The automatic image category annotation method based on density peak clustering algorithm according to claim 1, characterized in that, Calculate the reconstruction loss between the input unlabeled image data and the reconstructed image data. The specific expression is: Among them, n represents the size of the input unlabeled image dataset, and X i represents the input unlabeled image data sample, represents X i is the reconstructed image data obtained by passing through the convolutional encoder module and the convolutional decoder module. φ represents the parameters of the convolutional encoder module, θ represents the parameters of the convolutional decoder module, and f φ represents the mapping from the input unlabeled image data to the feature vector implemented by the convolutional encoder module, and g θ represents the mapping from the feature vector to the reconstructed image data implemented by the convolutional decoder module, and L rec represents the reconstruction loss of the entire convolutional autoencoder model.

4. The automatic image category annotation method based on density peak clustering algorithm according to claim 1, characterized in that, Using the density peak clustering method to calculate the local density and the distance to the high-density point of the feature vector points in the set of low-dimensional feature vectors. The specific method includes: Calculate the distance from each feature vector point to its k nearest neighbors, and calculate the mean μ and standard deviation σ of the k-nearest neighbor distances; where d(x, x i ) represents the Euclidean distance from the feature vector point x to its k-nearest neighbor x i ; According to the Chauvenet's criterion, an upper limit θ = μ + 3σ is calculated. The k-nearest neighbor distances are traversed, and the neighbors greater than the upper limit θ are removed to obtain a new k-nearest neighbor set. The local density ρ of the data point is calculated according to the updated k-nearest neighbor set as follows: where the updated k-nearest neighbor set is AKNN = {x j | d(x, x j ) ≤ d(x, x k ) ∧ d(x, x j ) ≤ θ = μ + 3σ}, x represents the feature vector point, x j is an object in the k-nearest neighbors of x, x j ∈ AKNN means that x j belongs to the k-nearest neighbors of x, d(x, x j ) represents the Euclidean distance between x and x j , and the symbol ^ represents conditional AND; The distance δ between the eigenvector point and the high-density point is: where ρ i represents the local density of point i, D represents the set of all eigenvector points, and d(x i , x j ) represents the Euclidean distance between x i and x j two points.

5. An automatic image category annotation device based on the density peak clustering algorithm, characterized in that, it includes: An image acquisition unit for acquiring an unlabeled image data set of the category to be annotated; A convolutional autoencoder model training unit for inputting the unlabeled image data set into the convolutional autoencoder model for training. The specific process includes: using the convolutional encoder module in the convolutional autoencoder model to reduce and compress the input unlabeled image data into a low-dimensional feature vector, and then using the convolutional decoder module to restore the low-dimensional feature vector to obtain the reconstructed image data. Calculate the reconstruction loss between the input unlabeled image data and the reconstructed image data. When the reconstruction loss is less than a given threshold, it is determined that the training of the convolutional autoencoder model stops; A low-dimensional feature vector set acquisition unit for retaining the convolutional encoder module in the trained convolutional autoencoder model and using the convolutional encoder module to obtain the low-dimensional feature vector set of the unlabeled image data set; A convolutional encoder-clustering joint training module training unit for inputting the low-dimensional feature vector set into the convolutional encoder-clustering joint training module for joint training. The specific process includes: using the density peak clustering method to calculate the local density and the distance from the high-density point of the feature vector points in the low-dimensional feature vector set, and multiplying the local density and the distance from the high-density point of the feature vector points to obtain the gamma value of the feature vector; descendingly sort the gamma values of all feature vectors in the low-dimensional feature vector set, and select the first m feature vector points as candidate cluster centers to obtain a candidate cluster center set; calculate the distances from the remaining feature vector points to each candidate cluster center to obtain an m-dimensional distance vector; take the reciprocal of each component in the m-dimensional distance vector and normalize it to obtain an m-dimensional class assignment confidence vector, Take the column where the component with the largest value in the m-dimensional class assignment confidence vector is located as the true class label of the feature vector point to obtain the clustering result of the class assignment confidence vector. Use the true class label to train the convolutional encoder module with labels; update the confidence vector matrix of the remaining feature vector points with the trained convolutional encoder module until the KL divergence value between the clustering results of two consecutive rounds is less than a given threshold, and the training ends; An annotated image data set unit for using the clustering result obtained after the training ends as the final clustering result, and annotating the input unlabeled image with the class division result to obtain the final annotated image data set.

6. The automatic image category annotation device based on the density peak clustering algorithm according to claim 5, characterized in that, Before inputting the unlabeled image data set into the convolutional autoencoder model for training, the convolutional autoencoder model training unit first performs data augmentation on the input unlabeled image data and adds random Gaussian noise.

7. An automatic image category annotation device based on the density peak clustering algorithm, characterized in that, it includes: A processor; and a memory, wherein the memory stores a computer-executable program, and when the computer-executable program is executed by the processor, the method for automatically annotating image categories based on the density peak clustering algorithm according to any one of claims 1-4 is executed.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the method for automatically annotating image categories based on the density peak clustering algorithm according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Driving mechanism.

    US940095A

  • Bearing fault diagnosis based on pseudo-tag semi-supervised kernel local Fisher discriminant analysis

    CN109582003A

  • Clustering autoencoder

    US20220129758A1