Image deduplication methods, apparatuses, electronic devices and storage media
By extracting image features and adjusting hierarchical clustering, the cluster structure is dynamically optimized, solving the problems of low efficiency and poor robustness in image deduplication in existing technologies. This achieves efficient and robust image deduplication, adapts to changes in data distribution, and supports an automated process for training deep learning models.
Patent Information
- Application Number
- CN202510948357.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-10
AI Technical Summary
In the existing technology, existing image deduplication methods cannot effectively solve the problem of the widespread use of image acquisition equipment. Existing image deduplication methods cannot efficiently remove duplicate or similar images in image annotation, model training, data storage management and other stages, resulting in waste of storage resources and increased processing costs.
By extracting image features and performing preliminary clustering, hierarchical clustering adjustments are made based on clustering quality evaluation indicators to dynamically optimize the cluster structure, achieving highly robust and efficient image deduplication, reducing manual intervention, and balancing computational complexity and deduplication accuracy.
It enables efficient removal of redundant images from massive image data, reduces storage resource consumption and processing costs, improves the robustness and efficiency of image deduplication, adapts to changes in data distribution, and supports automated processes for deep learning model training.
Smart Images

Figure CN120431355B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image deduplication method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the widespread adoption of image acquisition devices, massive amounts of image data have accumulated in internet platforms, social media, enterprise databases, and various smart terminals. This data contains a large number of repetitive or highly similar images, such as the same object photographed from different angles or the same content uploaded multiple times. This type of repetitive data not only consumes significant storage resources in image annotation, model training, and data storage management, but also increases data processing costs.
[0003] To address these issues, researchers have proposed image deduplication methods, which currently fall into two main categories: hash comparison and feature matching. Hash comparison quickly matches similar images using hash values, but its robustness is poor. When images undergo slight geometric deformations such as rotation or cropping, the hash values can change significantly, leading to semantically similar images being misclassified as dissimilar, resulting in a high false negative rate. Feature matching calculates similarity based on handcrafted features. This method is computationally complex and inefficient, making it difficult to meet real-time requirements. Furthermore, handcrafted features cannot fully represent the deep semantic information of images, and its performance degrades significantly in cross-domain data (such as the same scene under different lighting conditions). Summary of the Invention
[0004] This invention provides an image deduplication method, apparatus, electronic device, and storage medium to solve the problems of poor robustness, high computational complexity, and low efficiency in existing image deduplication methods. It achieves highly robust and efficient image deduplication, dynamically optimizes the cluster structure, and balances computational complexity and deduplication accuracy to solve the redundancy problem in massive image data.
[0005] This invention provides an image deduplication method, comprising:
[0006] Determine the set of images to be deduplicated, and extract the image features of each image in the set;
[0007] Based on the image features of each image, the image set is clustered to obtain multiple initial clusters;
[0008] Based on the clustering quality assessment index, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set; the clustering quality assessment index is used to reflect the cluster quality of each initial cluster.
[0009] According to an image deduplication method provided by the present invention, the step of performing hierarchical clustering adjustment on each initial cluster based on a clustering quality evaluation index to obtain the deduplication result corresponding to the image set includes:
[0010] Based on the clustering quality assessment indicators, including intra-cluster compactness, inter-cluster separation, and cluster profile coefficient, an adjustment strategy is determined.
[0011] Based on the adjustment strategy, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set.
[0012] According to an image deduplication method provided by the present invention, the adjustment strategy includes splitting and merging; the intra-cluster compactness is characterized by the cluster diameter, and the inter-cluster separation is characterized by the inter-cluster distance; the step of determining the adjustment strategy based on the intra-cluster compactness, inter-cluster separation, and cluster contour coefficient in the clustering quality evaluation index includes:
[0013] If the cluster profile coefficient of any initial cluster is less than the loose threshold and the cluster diameter of any initial cluster is greater than the diameter threshold, the adjustment strategy for any initial cluster is determined to be splitting.
[0014] If the inter-cluster distance between any initial cluster and any adjacent initial cluster is less than a compactness threshold, and the difference between the cluster profile coefficient of any initial cluster and the cluster profile coefficient of any adjacent initial cluster is less than a difference threshold, then the adjustment strategy for any initial cluster and any adjacent initial cluster is determined to be merging.
[0015] According to an image deduplication method provided by the present invention, the step of performing hierarchical clustering adjustment on each initial cluster based on the adjustment strategy to obtain the deduplication result corresponding to the image set includes:
[0016] Each initial cluster with the adjustment strategy of splitting is split, and two adjacent initial clusters with the adjustment strategy of merging are merged to obtain multiple candidate clusters;
[0017] The candidate clusters are updated to the initial clusters, and the updated initial clusters are adjusted hierarchically until the optimization termination condition is met. The optimization termination condition includes that the cluster profile coefficients of the initial clusters are all greater than or equal to the loose threshold and less than or equal to the compact threshold.
[0018] The initial clusters that satisfy the optimization termination condition are taken as target clusters, and the deduplication result corresponding to the image set is determined based on the multiple target clusters.
[0019] According to an image deduplication method provided by the present invention, the step of performing hierarchical clustering adjustment on each initial cluster based on a clustering quality evaluation index to obtain the deduplication result corresponding to the image set includes:
[0020] Based on the clustering quality evaluation index, hierarchical clustering adjustment is performed on each initial cluster to obtain multiple target clusters;
[0021] The target images in each target cluster are determined based on the distance between each image in each target cluster and the cluster center.
[0022] Based on the target images in each target cluster and the cluster labels of the target images, the deduplication result corresponding to the image set is determined.
[0023] According to an image deduplication method provided by the present invention, the step of extracting image features of each image in the image set includes:
[0024] Each image in the image set is input into the feature extraction model, which extracts features from each image to obtain the image features of each image output by the feature extraction model.
[0025] The feature extraction model is trained based on the feature similarity between sample image features in positive samples and the similarity between sample image features in negative samples; the positive samples include different sample augmented images of the same sample image, and the negative samples include sample augmented images of different sample images.
[0026] According to an image deduplication method provided by the present invention, the image set is clustered based on the image features of each image to obtain multiple initial clusters, including:
[0027] Determine the initial range of cluster values;
[0028] Based on the elbow algorithm, a candidate cluster value range is determined from the initial cluster value range;
[0029] Based on the silhouette coefficient method, the optimal initial cluster value is determined from the range of candidate cluster values;
[0030] Based on the image features of each image, the image set is clustered according to the optimal initial cluster value to obtain multiple initial clusters.
[0031] The present invention also provides an image deduplication device, comprising:
[0032] An extraction unit is used to determine the set of images to be deduplicated and to extract the image features of each image in the set of images;
[0033] A clustering unit is used to cluster the image set based on the image features of each image to obtain multiple initial clusters;
[0034] The deduplication unit is used to perform hierarchical clustering adjustment on each initial cluster based on the clustering quality evaluation index to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the image deduplication method as described above.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image deduplication method as described above.
[0037] The image deduplication method, apparatus, electronic device, and storage medium provided by this invention extract image features from each image in a set of images to be deduplicated; based on the image features of each image, the image set is clustered to obtain multiple initial clusters; based on a clustering quality evaluation index, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster, overcoming the shortcomings of traditional deduplication methods such as low efficiency, high computational complexity, and poor robustness. By performing preliminary clustering first, then evaluating cluster quality, and automatically triggering hierarchical clustering adjustment, a loop process for dynamically optimizing the cluster structure is achieved, which effectively solves the problems of overfitting and underfitting, avoids the imbalance of cluster structure, reduces manual intervention, and achieves highly robust and efficient image deduplication, effectively solving the redundancy problem of massive image data. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the image deduplication method provided by the present invention;
[0040] Figure 2 This is a flowchart of the node process for adjusting the hierarchical clustering provided by the present invention;
[0041] Figure 3 This is a schematic diagram of the image deduplication device provided by the present invention;
[0042] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0044] With the widespread adoption of image acquisition devices, massive amounts of image data have accumulated in internet platforms, social media, enterprise databases, and various smart terminals. This data contains a large number of duplicate or highly similar images, such as the same object photographed from different angles, the same content uploaded multiple times, and slightly edited or compressed image copies. This type of duplicate data not only consumes significant storage resources in image annotation, model training, and data storage management, but also significantly increases data processing costs. For example, in deep learning model training, duplicate images can lead to skewed training sample distribution, reducing the model's generalization ability; in data storage systems, redundant data directly causes wasted storage space and increased maintenance costs. However, traditional manual screening methods for image deduplication are inefficient and cannot effectively solve the redundancy problem of massive image data.
[0045] Based on this, automated image deduplication methods have emerged, and the current automated image deduplication technologies mainly fall into the following categories:
[0046] First, hash comparison: This method, represented by Perceptual Hash Algorithm (PHA) and Average Hash Algorithm (AHA), achieves fast similarity matching by converting images into fixed-length hash values. This method has low computational complexity and is suitable for initial screening of large-scale data. However, its core drawback lies in its insufficient robustness to image transformations: when images undergo rotation, cropping, brightness adjustment, or slight geometric deformation, the hash value may change significantly, leading to semantically similar but pixel-level differences being misclassified as dissimilar, resulting in a high false negative rate.
[0047] Second, the feature matching method: This method extracts handcrafted features such as SIFT (Scale-invariant feature transform) and SURF (Speeded Up Robust Features), and calculates image similarity through feature point matching. This method relies on high-dimensional feature extraction and complex distance calculations, resulting in high computational complexity and low efficiency, making it difficult to meet real-time requirements. Furthermore, handcrafted features cannot fully represent the deep semantic information of images, and performance degrades significantly in cross-domain data (such as the same scene under different lighting conditions).
[0048] Third, fixed clustering: This method uses traditional clustering algorithms such as K-means to divide the cluster structure with a fixed K value. It achieves deduplication by reducing intra-cluster similarity. However, this method has two major drawbacks: On the one hand, it requires manual intervention to pre-set the number of clusters K, and lacks the ability to dynamically adapt to data distribution; on the other hand, a fixed K value cannot cope with the growth of data scale or changes in category distribution, which can easily lead to an imbalance in cluster structure. For example, excessive merging of similar images or excessive splitting of the same category makes it difficult to balance deduplication accuracy and recall.
[0049] Therefore, there is an urgent need for a robust, efficient, and adaptive image deduplication technology that can dynamically optimize the cluster structure without human intervention, while balancing computational complexity and deduplication accuracy, in order to solve the redundancy problem in massive image data.
[0050] To address this, the present invention provides an image deduplication method that aims to overcome the shortcomings of existing solutions, achieve highly robust and efficient image deduplication, and dynamically optimize the cluster structure through an adaptive clustering mechanism to reduce manual intervention and achieve a balance between deduplication accuracy and efficiency. Figure 1 This is a flowchart illustrating the image deduplication method provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0051] Step 110: Determine the set of images to be deduplicated and extract the image features of each image in the set;
[0052] Step 120: Based on the image features of each image, cluster the image set to obtain multiple initial clusters;
[0053] Step 130: Based on the clustering quality assessment index, perform hierarchical clustering adjustment on each initial cluster to obtain the deduplication result corresponding to the image set; the clustering quality assessment index is used to reflect the cluster quality of each initial cluster.
[0054] Specifically, considering the shortcomings of current image deduplication methods in terms of efficiency, robustness, and adaptability—namely, hash comparison methods sacrifice robustness for speed, feature matching methods trade computational overhead for accuracy, and fixed clustering methods are unable to adapt to dynamic data environments due to static parameter limitations—especially in deep learning model training scenarios, during the pre-processing image annotation and deep learning training set cleaning processes, current image deduplication methods struggle to balance deduplication efficiency and semantic consistency. This results in a significant need for manual verification during data preprocessing, severely hindering the implementation of automated processes.
[0055] In view of this, in this embodiment of the invention, it is proposed that preliminary clustering be performed first, and then the quality of the cluster structure be evaluated based on the data distribution. Based on this, hierarchical clustering adjustment is automatically triggered to dynamically optimize the cluster structure, while solving the problems of overfitting and underfitting, avoiding cluster structure imbalance, and making the clustering results adapt to the data distribution. This breaks through the limitation of traditional clustering algorithms that require a fixed number of clusters, and achieves a balance between deduplication accuracy and efficiency.
[0056] In detail, in practical applications, before performing image deduplication, it is necessary to first determine the set of images to be deduplicated, i.e., the set of images to be deduplicated. Here, the set of images to be deduplicated can be one or more. It can be a collection of images from the same device / platform / enterprise database / social media account, or it can be a collection of images from different devices / platforms / enterprise databases / social media accounts. This embodiment of the invention does not specifically limit this. The images can be from a single field or category, or they can be from multiple fields or categories. This embodiment of the invention does not specifically limit this.
[0057] Once the set of images to be deduplicated is determined, in this embodiment of the invention, feature extraction can be performed on this set of images to extract the image features of each image for subsequent clustering to achieve deduplication. Specifically, feature extraction can be performed on each image in the image set to extract information that can represent image characteristics, such as uniqueness (distinguishing different images), semantic similarity (identifying duplicate content), and visual invariance (anti-interference ability), thereby obtaining the image features of each image.
[0058] Here, feature extraction targeting image characteristics can improve the robustness and discriminativeness of the extracted image features. Improved robustness enhances the ability to resist interference, ensuring the stability of image features under image transformations (such as rotation, scaling, and cropping). Improved discriminativeness helps increase recognition accuracy, enabling image features to distinguish semantically similar but pixel-level different images (such as the same object photographed from different angles), thereby avoiding missed detections.
[0059] The above feature extraction process can be implemented based on deep learning feature extraction methods. That is, global features of the image are extracted by a pre-trained feature extraction model, such as ResNet (Residual Network) 50. Then, the last classification layer of the model is removed, and the output of the second to last layer is taken as the feature vector. The feature vector is then normalized to obtain the image features of each image. The image features obtained in this way have strong semantic discriminativeness, are suitable for cross-domain images, and have low computational complexity and high efficiency when performing subsequent clustering processing based on these image features.
[0060] Feature extraction can also be achieved through local and global feature extraction methods. First, SIFT or ORB (Oriented Fast and Rotated BRIEF) is used to extract local keypoints and descriptors. Then, BoVW (Bag of Visual Words) or VLAD (Vector of Locally Aggregated Descriptors) is used to aggregate the local features into a global vector. Next, PCA (Principal Components Analysis) is used to reduce the dimensionality of the global vector to obtain image features. The image features obtained in this way have strong robustness and interpretability, are not sensitive to geometric transformations, and allow for visualization of local features.
[0061] Feature extraction can also be achieved through a self-supervised learning-based feature extraction method. This involves first using a self-supervised model, such as SimCLR (Simple Framework for Contrastive Learning of Visual Representations), pre-trained on large-scale unlabeled data, and then using the trained model to extract features from each image in the image set. The intermediate layer features are then extracted as image representations, resulting in image features with strong adaptability and generalization ability.
[0062] Feature extraction can also be achieved in other ways, such as hybrid feature fusion, which combines deep learning features with handcrafted features (such as color histograms and LBP textures) to achieve a balance between semantics and low-level information and improve deduplication accuracy; or, for example, high-level semantic feature extraction based on residual networks. This embodiment of the invention does not specifically limit this.
[0063] After obtaining the image features of each image in the image set, this embodiment of the invention can cluster the image set based on these image features to divide the images in the image set into multiple clusters through preliminary clustering, thereby obtaining multiple initial clusters. The clustering algorithm used in this preliminary clustering process can be the traditional K-means clustering algorithm. In this preliminary clustering process, the K value can be predetermined based on the elbow algorithm, silhouette coefficient method, etc., that is, determining the optimal preliminary clustering K value. Based on this, preliminary clustering can improve the clustering quality, balance intra-cluster compactness and inter-cluster separation, reduce subjectivity, optimize computational efficiency, avoid unnecessary clustering calculations, and reduce computational overhead. Furthermore, determining the optimal K value can significantly improve the performance and reliability of the clustering algorithm in image deduplication tasks, thereby helping to optimize the deduplication accuracy and efficiency of subsequent image deduplication tasks.
[0064] Furthermore, considering that a single clustering based on the optimal K value in the initial clustering may not guarantee the cluster structure and may have overfitting and underfitting problems, leading to an imbalance in the cluster structure, this embodiment of the invention proposes to perform a quality assessment on the clusters obtained from the initial clustering to determine the cluster quality of each initial cluster, and based on this, to optimize and adjust the results of the initial clustering, i.e., each initial cluster, to obtain the final deduplication result.
[0065] Specifically, this can begin by determining preliminary clustering quality assessment metrics, which reflect the cluster quality of each initial cluster. These metrics may include intra-cluster compactness, inter-cluster separation, cluster profile coefficient, Davidson-Bolding index (a measure of the ratio of intra-cluster distance to inter-cluster distance, with smaller values indicating better clustering), Calinski-Harabasz index (a measure of the ratio of inter-cluster dispersion to intra-cluster dispersion, with larger values indicating better clustering), and Dunn index (a measure of the ratio of minimum inter-cluster distance to maximum intra-cluster diameter, with larger values indicating better clustering). Next, based on these clustering quality assessment metrics, hierarchical clustering adjustments can be made to each initial cluster. By simulating the splitting and merging process of hierarchical clustering, the cluster structure can be adjusted, such as splitting clusters with excessively low intra-cluster compactness or merging clusters with excessively low inter-cluster separation, thereby obtaining optimized clustering results. The final deduplication result can then be determined based on these clustering results. The deduplication result here can include the deduplicated image set and the cluster labels of each image in the image set. The cluster labels can be the cluster number, number of splits or merges, etc. of the corresponding image.
[0066] It is worth noting that the hierarchical clustering adjustment process in this embodiment of the invention can be executed once, i.e., in one step, to improve deduplication efficiency and save computational overhead, or it can be executed repeatedly. That is, after the first step of adjustment, the clusters obtained by adjustment are used as the initial clusters, and the hierarchical clustering adjustment process is executed again to further optimize the cluster structure in a manner that simulates hierarchical clustering, so as to ensure deduplication accuracy. Alternatively, the number of executions can be automatically matched according to specific circumstances, i.e., an optimization termination condition can be set, and the optimization will automatically stop when the condition is met, so as to balance efficiency and deduplication accuracy. This embodiment of the invention does not make specific limitations on this.
[0067] In this embodiment of the invention, based on the initial clustering, the cluster structure quality is quantitatively evaluated, and the cluster structure optimization mechanism is automatically triggered. The hierarchical clustering adjustment solves the underfitting and overfitting problems that may exist in the initial clustering, so that the final clustering result adapts to the data distribution. This breaks through the dilemma of the traditional K-means clustering algorithm having a fixed number of clusters, which cannot cope with the growth of data scale or changes in category distribution, and is prone to cluster structure imbalance. It achieves a balance between deduplication efficiency and accuracy.
[0068] The image deduplication method provided by this invention extracts image features from each image in a set of images to be deduplicated; based on the image features of each image, the image set is clustered to obtain multiple initial clusters; based on the clustering quality evaluation index, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster, overcoming the shortcomings of traditional deduplication methods such as low efficiency, high computational complexity, and poor robustness. By first performing preliminary clustering, then evaluating the cluster quality, and automatically triggering hierarchical clustering adjustment, a loop process of dynamically optimizing the cluster structure is achieved, which effectively solves the problems of overfitting and underfitting, avoids the imbalance of the cluster structure, reduces manual intervention, and achieves highly robust and efficient image deduplication, effectively solving the redundancy problem of massive image data.
[0069] In addition, it should be noted that the embodiments of the present invention can achieve efficient and robust image deduplication without sacrificing robustness for speed or computational overhead for accuracy. In deep learning model training scenarios, such as the pre-image annotation process and the cleaning process of deep learning training sets, it can better balance deduplication efficiency and semantic consistency, thereby facilitating the implementation of automated processes for deep learning model training.
[0070] Based on the above embodiments, step 130 includes:
[0071] Based on the cluster quality assessment indicators, including intra-cluster compactness, inter-cluster separation, and cluster profile coefficient, an adjustment strategy is determined.
[0072] Based on the adjustment strategy, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set.
[0073] Specifically, the process of adjusting each initial cluster according to the clustering quality evaluation index to obtain the deduplication result corresponding to the image set may include the following steps:
[0074] First, cluster quality evaluation indicators can be determined to assess the cluster quality of the initial clusters. Preferably, in this embodiment of the invention, intra-cluster compactness, inter-cluster separation, and cluster profile coefficient are used as evaluation indicators to assess the cluster quality of the corresponding initial clusters, thereby facilitating subsequent adjustments to their cluster structure.
[0075] Here, intra-cluster compactness can be characterized by the cluster diameter corresponding to the initial cluster. The cluster diameter is the maximum distance between all data points (images) within the cluster, which measures the dispersion of data points within the cluster. The smaller the cluster diameter, the more compact the data points within the cluster; conversely, the larger the cluster diameter, the more dispersed the data points within the cluster. Intra-cluster compactness can also be represented by the average cluster distance, i.e., the average distance from the data points within the cluster to the cluster center; it can also be represented by the intra-cluster variance, i.e., the mean of the squared distances from the data points within the cluster to the cluster center; or it can be represented by other indicators. This embodiment of the invention does not specifically limit the specific methods used.
[0076] Inter-cluster separation can be characterized by inter-cluster distance, which is the distance between the cluster centers of two adjacent clusters. It measures the degree of separation between clusters; the larger the inter-cluster distance, the more obvious the differences between clusters, and the higher the separation. Inter-cluster separation can also be represented by the average inter-cluster distance, which is the average distance between all data points in one initial cluster and all other initial clusters; it can also be represented by the Dunn exponent. This embodiment of the invention does not specifically limit the representation of this method.
[0077] The cluster profile coefficient is a comprehensive index that combines intra-cluster cohesion and inter-cluster separation. It is used to evaluate the reasonableness of a single data point or the overall clustering result. A cluster profile coefficient as close to 1 as possible indicates a better clustering effect.
[0078] Next, the adjustment strategy for the corresponding initial cluster can be determined based on the intra-cluster compactness, inter-cluster separation, and cluster profile coefficient. That is, based on these indicators, it is assessed whether the cluster structure of the corresponding initial cluster needs to be adjusted, and how to make the adjustment.
[0079] For example, if the distribution of data points within an initial cluster is determined to be scattered based on the cluster compactness, and / or the distribution of data points in the initial cluster is deemed unreasonable based on the cluster profile coefficient, then it can be determined that the cluster structure of the initial cluster needs to be adjusted. In other words, the initial cluster can be split into two sub-clusters.
[0080] For example, if the distance between the cluster centers of an initial cluster and its neighboring initial clusters is small based on the inter-cluster separation, and the separation degree of the adjacent clusters is low, and / or the data point distribution of the two adjacent initial clusters is not reasonable based on the cluster profile coefficient, then it can be determined that the cluster structure of the two adjacent initial clusters needs to be adjusted, that is, the two adjacent initial clusters can be merged into one cluster.
[0081] Then, according to the adjustment strategy, hierarchical clustering can be performed on each initial cluster. By simulating the splitting and merging process of hierarchical clustering, the cluster structure can be adjusted to obtain the optimized clustering result. The final deduplication result can be determined based on this clustering result. The deduplication result here can include the deduplicated image set and the cluster labels of each image in the image set. The cluster labels can be the cluster number of the corresponding image, the number of splits or merges, etc.
[0082] Based on the above embodiments, intra-cluster compactness is characterized by cluster diameter, and inter-cluster separation is characterized by inter-cluster distance; adjustment strategies include splitting and merging.
[0083] Based on the clustering quality assessment metrics of intra-cluster compactness, inter-cluster separation, and cluster profile coefficient, adjustment strategies are determined, including:
[0084] If the cluster profile coefficient of any initial cluster is less than the loose threshold and the cluster diameter of any initial cluster is greater than the diameter threshold, the adjustment strategy for any initial cluster is determined to be splitting.
[0085] If the inter-cluster distance between any initial cluster and any adjacent initial cluster is less than the compactness threshold, and the difference between the cluster profile coefficient of any initial cluster and the cluster profile coefficient of any adjacent initial cluster is less than the difference threshold, then the adjustment strategy for any initial cluster and any adjacent initial cluster is determined to be merging.
[0086] Specifically, the process of determining adjustment strategies based on cluster quality assessment metrics such as intra-cluster compactness, inter-cluster separation, and cluster profile coefficient can include the following two categories:
[0087] First, when the cluster profile coefficient of any initial cluster is less than the loose threshold, and the intra-cluster compactness of the initial cluster, i.e. the cluster diameter, is greater than the diameter threshold, it means that the data points within the initial cluster are scattered and not reasonable. In this case, it can be determined that the cluster structure of the initial cluster needs to be adjusted, and the adjustment strategy is splitting.
[0088] Secondly, when the distance between any initial cluster and any adjacent initial cluster is less than the compactness threshold, and the difference in the cluster profile coefficients of these two initial clusters is less than the difference threshold, it means that the distance between the cluster centers of these two initial clusters is too close, the separation degree of the two clusters is low, and the distribution of data points within the clusters is not reasonable. In this case, it can be determined that the cluster structure of these two adjacent initial clusters needs to be adjusted, and the adjustment strategy is to merge them.
[0089] Among them, the loose threshold, diameter threshold, compact threshold and difference threshold can all be set according to the actual situation and actual needs. For example, the loose threshold can be 0.2, the compact threshold can be 0.7, and the diameter threshold can be determined according to the global average distance (the average of the average distances of all clusters), such as twice the global average distance.
[0090] The following uses specific numerical examples to illustrate the process of determining the adjustment strategy:
[0091] For the initial cluster Calculate its cluster profile coefficient and inter-cluster distance , ,in for The average distance between each data point and other data points. for The average distance between each data point and all data points in the nearest other initial cluster. ,in for The cluster center, for Adjacent initial clusters The cluster center.
[0092] like And cluster diameter , If the diameter threshold is used, then the determination is made for... The adjustment strategy is to split, dividing into two sub-clusters. For example, to... The "animal" cluster splits into the "cat" cluster and the "dog" cluster.
[0093] like ,and , If the difference threshold is used, then the determination is made for... and The adjustment strategy is to merge, merging two clusters. For example, to... (The "Husky" cluster) and (The "Alaska" cluster) merged.
[0094] Based on the above embodiments, and based on the adjustment strategy, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set, including:
[0095] Each initial cluster with the adjustment strategy of splitting is split, and two adjacent initial clusters with the adjustment strategy of merging are merged to obtain multiple candidate clusters;
[0096] Multiple candidate clusters are updated to multiple initial clusters, and hierarchical clustering is performed on the updated initial clusters until the optimization termination condition is met. The optimization termination condition includes that the cluster profile coefficients of multiple initial clusters are all greater than or equal to the loose threshold and less than or equal to the compact threshold.
[0097] Multiple initial clusters that meet the optimization termination conditions are used as target clusters, and the deduplication result corresponding to the image set is determined based on multiple target clusters.
[0098] Specifically, after determining the adjustment strategy corresponding to each initial cluster, hierarchical clustering adjustment can be performed on each initial cluster according to this adjustment strategy to obtain the deduplication result. This process can specifically include:
[0099] First, the cluster structure of the corresponding initial clusters can be adjusted according to the adjustment strategy. That is, the initial clusters with the adjustment strategy of splitting are split, and the two adjacent initial clusters with the adjustment strategy of merging are merged. In this way, multiple clusters with preliminary optimization can be obtained, that is, multiple candidate clusters.
[0100] Next, these candidate clusters can be used as initial clusters for hierarchical clustering adjustments to further optimize the cluster structure of the initially optimized clusters. The specific optimization process involves determining the adjustment strategy for each updated initial cluster and performing hierarchical clustering adjustments according to this strategy. The determination of the adjustment strategy here is the same as that for the initial clusters before the update, as detailed above and will not be repeated here. This cluster structure optimization process is repeated until the optimization termination condition is met. The optimization termination condition here includes that the cluster profile coefficients of all initial clusters are greater than or equal to the loose threshold and less than or equal to the compact threshold; that is, the cluster profile coefficients of all initial clusters satisfy... .
[0101] Then, the final deduplication result can be determined based on the multiple initial clusters that meet the optimization termination condition. Specifically, the multiple initial clusters that meet the optimization termination condition can be used as target clusters, which are the final clustering results. Each target cluster is a set of duplicate images, containing at least one image. Therefore, the most representative image can be selected as the image representative of this set of duplicate images, i.e., the target image. At the same time, the other images in this cluster are marked as duplicate images. Based on the target image of each target cluster, the deduplicated image set can be determined.
[0102] Furthermore, based on the deduplicated image set and the cluster label of each target image, the final deduplication result can be determined. Here, the cluster label is the cluster number of the cluster to which the target image belongs.
[0103] Based on the above embodiments, step 130 includes:
[0104] Based on the clustering quality assessment index, hierarchical clustering adjustment is performed on each initial cluster to obtain multiple target clusters;
[0105] The target images in each target cluster are determined based on the distance between each image in each target cluster and the cluster center.
[0106] Based on the target images in each target cluster and the cluster labels of the target images, the deduplication result corresponding to the image set is determined.
[0107] Specifically, the process of adjusting the initial clusters hierarchically based on clustering quality evaluation metrics to obtain the deduplication result for the image set includes:
[0108] After hierarchical clustering adjustment, multiple adjusted clusters can be obtained, which are referred to here as target clusters. The hierarchical clustering adjustment process for each initial cluster has been described in detail above and will not be repeated here. Next, the most representative image can be selected from each target cluster as the target image. Specifically, since the target cluster is the final clustering result, each target cluster is a set of repeated images containing at least one image. Therefore, in each target cluster, the distance of each image from the cluster center can be determined. The closer the image is to the cluster center, the more representative it is. Therefore, the image closest to the cluster center can be selected as the target image, and the other images are marked as repeated images.
[0109] After this, the deduplicated image set can be determined based on the target images of each target cluster. Further, the final deduplication result can be determined based on the deduplicated image set and the cluster label of each target image in that image set. Here, the cluster label is the cluster number of the cluster to which the target image belongs.
[0110] Figure 2 This is a flowchart of the node process for hierarchical clustering adjustment provided by the present invention, as follows: Figure 2 As shown, after obtaining multiple initial clusters through preliminary clustering, the cluster quality of each initial cluster can be evaluated before hierarchical clustering adjustments. Specifically, this can be based on the intra-cluster compactness, inter-cluster separation, and cluster profile coefficient of each initial cluster—i.e., cluster quality evaluation indicators—to reflect cluster quality, and adjustment strategies can be determined based on these indicators. Adjustment strategies include splitting and merging. Intra-cluster compactness is characterized by cluster diameter, and inter-cluster separation is characterized by inter-cluster distance. That is, if the cluster profile coefficient of any initial cluster is less than the looseness threshold, and the cluster diameter of any initial cluster is greater than the diameter threshold, then the cluster quality is considered optimized. If the initial cluster meets the splitting condition, the adjustment strategy for any initial cluster is determined to be splitting. This can also be understood as the initial cluster meeting the splitting condition and being split. Correspondingly, if the inter-cluster distance between any initial cluster and any adjacent initial cluster is less than the compactness threshold, and the difference between the cluster profile coefficient of any initial cluster and the cluster profile coefficient of any adjacent initial cluster is less than the difference threshold, the adjustment strategy for any initial cluster and any adjacent initial cluster is determined to be merging. This can also be understood as the initial cluster meeting the merging condition and being merged.
[0111] After splitting and merging, multiple candidate clusters are obtained. Further optimization and adjustment of these clusters' structures are then required. Specifically, the candidate clusters are updated to initial clusters, and hierarchical clustering adjustments are performed on these updated initial clusters until the optimization termination condition is met. The optimization termination condition includes that the cluster contour coefficients of the multiple initial clusters are all greater than or equal to a loose threshold and less than or equal to a compact threshold. Finally, the initial clusters that meet the optimization termination condition can be used as target clusters. Based on the distance between each image in each target cluster and the cluster center, the image closest to the cluster center is selected as the target image. Based on the target images of each target cluster, the deduplicated image set can be determined. Further, based on the deduplicated image set and the cluster label of each target image in the image set, the final deduplication result can be determined. Here, the cluster label is the cluster number of the cluster to which the target image belongs.
[0112] Based on the above embodiments, step 110, extracting image features from each image in the image set, includes:
[0113] Each image in the image set is input into the feature extraction model, which extracts features from each image to obtain the image features of each image output by the feature extraction model.
[0114] The feature extraction model is trained based on the feature similarity between sample image features in positive samples and the similarity between sample image features in negative samples; positive samples include different sample augmented images of the same sample image, and negative samples include sample augmented images of different sample images.
[0115] Specifically, in step 110, the feature extraction process for the image set can be implemented through a feature extraction model. Specifically, each image in the image set can be input into the feature extraction model, and the feature extraction model can extract features from the input images to extract information that can represent image characteristics, such as uniqueness (distinguishing different images), semantic similarity (identifying duplicate content), and visual invariance (anti-interference ability), and encode them as features to obtain low-dimensional (2048-dimensional) image features with high semantic discriminative power output by the feature extraction model.
[0116] It is worth noting that, before performing feature extraction using the feature extraction model, a feature extraction model can be pre-trained to ensure the accuracy of feature extraction. Unlike traditional supervised and unsupervised learning, this embodiment of the invention considers that image deduplication tasks require features to have strong anti-interference and discriminative capabilities—that is, the ability to resist image transformations and identify semantically similar but pixel-level different images. Therefore, the similarity of semantic information represented by different sample images is used for model training to obtain a trained feature extraction model.
[0117] Specifically, during model training, a large number of sample images need to be collected first. Then, image enhancement can be applied to the same sample image using different enhancement techniques, such as rotation, cropping, scaling, and color dithering, resulting in multiple enhanced sample images. Subsequently, positive and negative samples can be constructed from these enhanced images. Positive samples are enhanced versions of the same corresponding sample image, while negative samples are enhanced versions of different corresponding sample images. Based on this principle, the positive and negative samples required for training can be constructed from the enhanced sample images obtained. Afterward, the initial model can be used to determine the sample image features of the enhanced images in the positive and negative samples, respectively.
[0118] The sample image features are obtained by the initial model from the feature extraction of the corresponding sample enhanced image; the initial model here can be built on a deep convolutional neural network model, such as ResNet50.
[0119] Furthermore, after determining the sample image features of the augmented images in the positive samples and the sample image features of the augmented images in the negative samples, in this embodiment of the invention, the contrast loss can be determined based on the feature similarity between the sample image features in the positive samples and the similarity between the sample image features in the negative samples. Based on this contrast loss, the initial model is trained to obtain the trained model, i.e., the feature extraction model. The contrast loss here can be measured using InfoNCE (Information Noise-Contrastive Estimation).
[0120] Compared to traditional methods that extract image features from different sample images and perform duplicate image detection based on these features, using the error between predicted and labeled values to drive model parameter updates, this invention selects the semantic similarity represented by different sample images for model training. By training the initial model through the feature similarity between sample image features in positive samples and the similarity between sample image features in negative samples, the initial model can fully learn the proximity relationship between sample image features corresponding to different sample augmented images, thus providing crucial assistance for subsequent clustering and deduplication processing.
[0121] Specifically, the initial model training objective is to maximize the feature similarity between the features of different enhanced images when different enhanced images constitute positive samples (i.e., the sample images corresponding to different enhanced images are the same), and conversely, to minimize the feature similarity between the features of different enhanced images when different enhanced images constitute negative samples (i.e., the sample images corresponding to different enhanced images are different). Therefore, when the feature similarity between the features of enhanced images in positive samples is high and the feature similarity between the features of enhanced images in negative samples is low, the contrast loss can be determined to be small. Conversely, when the feature similarity between the features of enhanced images in positive samples is low and / or the feature similarity between the features of enhanced images in negative samples is high, the contrast loss can be determined to be large.
[0122] In this embodiment of the invention, the semantic similarity of the same image on different enhanced images is used to train the model, which can not only improve the generalization ability of the model, but also improve the clustering accuracy and the reliability of the deduplication results.
[0123] Based on the above embodiments, step 120 includes:
[0124] Determine the initial range of cluster values;
[0125] Based on the elbow algorithm, the range of candidate cluster values is determined from the initial range of cluster values;
[0126] Based on the silhouette coefficient method, the optimal initial cluster value is determined from the range of candidate cluster values;
[0127] Based on the image features of each image, the image set is clustered according to the optimal initial cluster value to obtain multiple initial clusters.
[0128] Specifically, the process of clustering the image set based on the image features of each image to obtain multiple initial clusters may include:
[0129] First, determine the preliminary... The value search range, i.e. the initial cluster value range, can be [2, 100] or [2, 50], and can be set according to the number of images in the image set to be deduplicated, the deduplication accuracy, etc.
[0130] Then, for each of these ranges... Values, perform K-means clustering, and compute each Clustering error of values , or Sum of Squared Errors, represents the sum of squared Euclidean distances from all data points within a cluster to its cluster center.
[0131] Here, the formula for calculating clustering error is:
[0132]
[0133] In the formula, For clustering error, For the first Clusters, The number of clusters. for Data points / images in the image for The cluster center, for arrive The Euclidean distance.
[0134] Then, it can The x-axis represents the clustering values, and the y-axis represents the corresponding clustering errors. The curve, find the "elbow point", that is The point where the rate of descent slows significantly can be used to quickly narrow down the range. The range of values is used to determine the optimal value from the initial range of cluster values. The value can be narrowed based on the "elbow point" to obtain a smaller range of candidate cluster values.
[0135] Here, the elbow algorithm is used for range narrowing, which can quickly and accurately filter out candidates. Values should be considered to avoid subjective assumptions and reduce the cost of trial and error.
[0136] Furthermore, based on the range of candidate cluster values, the silhouette coefficient method can be used for further screening to select the optimal cluster. Value. Specifically, this could be for each value within the range of candidate cluster values. The values are clustered, and the average silhouette coefficient of each cluster is calculated. This average is obtained by averaging the silhouette coefficients of all data points within that cluster. Furthermore, the average silhouette coefficients of all clusters can be averaged again to obtain the average silhouette coefficient of each cluster. The overall profile score is the value of the clustering. The higher the overall profile score (closer to 1), the better the clustering effect (closer within clusters and more dispersed between clusters). Using the overall profile score to measure the clustering effect can avoid subjective judgment, reduce human bias, and adapt to complex data, that is, it can still give an effective evaluation for non-convex distribution or noisy data.
[0137] After this, the values of each candidate cluster can be directly compared within their respective ranges. The overall profile score is selected based on the highest overall profile score. Value, as the optimal The optimal initial cluster value is then determined. Subsequently, the image set can be clustered according to the image features of each image to divide the images in the image set into clusters with the optimal initial cluster value, thereby obtaining multiple initial clusters.
[0138] In this embodiment of the invention, by combining the elbow algorithm and the silhouette coefficient method to determine the optimal initial cluster values, and performing preliminary clustering based on these values, the reliability and efficiency of the preliminary clustering results can be significantly improved. Specifically, the elbow algorithm analyzes the variation of clustering error with... The downward trend of value changes allows for rapid identification of the value range of candidate clusters, which is particularly suitable for preliminary exploration of data distribution and capturing possible "elbow points," thus providing a reasonable basis for subsequent analysis. The value range; while the silhouette coefficient rule, based on this, quantifies the intra-cluster compactness and inter-cluster separation of each data point for all candidate values. The values are evaluated in detail to ensure that the final selection is optimal. This approach maximizes intra-cluster similarity while minimizing inter-cluster overlap, ultimately compensating for potential misjudgments caused by elbow point ambiguity or data noise in the elbow algorithm, thus improving objectivity and accuracy. This combined strategy not only avoids the limitations of single methods (such as the subjectivity of the elbow algorithm or the high computational cost of the contour coefficient method) but also optimizes computational resource allocation through a "coarse screening followed by fine adjustment" process. Especially in large-scale image data, combining this with sampling or Mini-Batch K-Means approximation algorithms can further reduce computational costs, ultimately achieving a dual improvement in clustering quality and efficiency. It can also contribute to optimizing the deduplication accuracy and efficiency of image deduplication tasks.
[0139] The image deduplication device provided by the present invention is described below. The image deduplication device described below can be referred to in correspondence with the image deduplication method described above.
[0140] Figure 3 This is a schematic diagram of the image deduplication device provided by the present invention, as shown below. Figure 3 As shown, the device includes:
[0141] Extraction unit 310 is used to determine the set of images to be deduplicated and extract the image features of each image in the set of images;
[0142] Clustering unit 320 is used to cluster the image set based on the image features of each image to obtain multiple initial clusters;
[0143] The deduplication unit 330 is used to perform hierarchical clustering adjustment on each initial cluster based on the clustering quality evaluation index to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster.
[0144] The image deduplication device provided by this invention extracts image features from each image in a set of images to be deduplicated; based on the image features of each image, the image set is clustered to obtain multiple initial clusters; based on the clustering quality evaluation index, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster, overcoming the shortcomings of traditional deduplication methods such as low efficiency, high computational complexity, and poor robustness. By first performing preliminary clustering, then evaluating the cluster quality, and automatically triggering hierarchical clustering adjustment, a loop process for dynamically optimizing the cluster structure is achieved, which effectively solves the problems of overfitting and underfitting, avoids the imbalance of the cluster structure, reduces manual intervention, and achieves highly robust and efficient image deduplication, effectively solving the redundancy problem of massive image data.
[0145] Based on the above embodiments, the deduplication unit 330 is used for:
[0146] Based on the clustering quality assessment indicators, including intra-cluster compactness, inter-cluster separation, and cluster profile coefficient, an adjustment strategy is determined.
[0147] Based on the adjustment strategy, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set.
[0148] Based on the above embodiments, the adjustment strategy includes splitting and merging; the intra-cluster compactness is characterized by the cluster diameter, and the inter-cluster separation is characterized by the inter-cluster distance;
[0149] Deduplication unit 330 is used for:
[0150] If the cluster profile coefficient of any initial cluster is less than the loose threshold and the cluster diameter of any initial cluster is greater than the diameter threshold, the adjustment strategy for any initial cluster is determined to be splitting.
[0151] If the inter-cluster distance between any initial cluster and any adjacent initial cluster is less than a compactness threshold, and the difference between the cluster profile coefficient of any initial cluster and the cluster profile coefficient of any adjacent initial cluster is less than a difference threshold, then the adjustment strategy for any initial cluster and any adjacent initial cluster is determined to be merging.
[0152] Based on the above embodiments, the deduplication unit 330 is used for:
[0153] Each initial cluster with the adjustment strategy of splitting is split, and two adjacent initial clusters with the adjustment strategy of merging are merged to obtain multiple candidate clusters;
[0154] The candidate clusters are updated to the initial clusters, and the updated initial clusters are adjusted hierarchically until the optimization termination condition is met. The optimization termination condition includes that the cluster profile coefficients of the initial clusters are all greater than or equal to the loose threshold and less than or equal to the compact threshold.
[0155] The initial clusters that satisfy the optimization termination condition are taken as target clusters, and the deduplication result corresponding to the image set is determined based on the multiple target clusters.
[0156] Based on the above embodiments, the deduplication unit 330 is used for:
[0157] Based on the clustering quality evaluation index, hierarchical clustering adjustment is performed on each initial cluster to obtain multiple target clusters;
[0158] The target images in each target cluster are determined based on the distance between each image in each target cluster and the cluster center.
[0159] Based on the target images in each target cluster and the cluster labels of the target images, the deduplication result corresponding to the image set is determined.
[0160] Based on the above embodiments, the extraction unit 310 is used for:
[0161] Each image in the image set is input into the feature extraction model, which extracts features from each image to obtain the image features of each image output by the feature extraction model.
[0162] The feature extraction model is trained based on the feature similarity between sample image features in positive samples and the similarity between sample image features in negative samples; the positive samples include different sample augmented images of the same sample image, and the negative samples include sample augmented images of different sample images.
[0163] Based on the above embodiments, the clustering unit 320 is used for:
[0164] Determine the initial range of cluster values;
[0165] Based on the elbow algorithm, a candidate cluster value range is determined from the initial cluster value range;
[0166] Based on the silhouette coefficient method, the optimal initial cluster value is determined from the range of candidate cluster values;
[0167] Based on the image features of each image, the image set is clustered according to the optimal initial cluster value to obtain multiple initial clusters.
[0168] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an image deduplication method. This method includes: determining a set of images to be deduplicated and extracting image features from each image in the set; clustering the image set based on the image features to obtain multiple initial clusters; and performing hierarchical clustering adjustment on each initial cluster based on a clustering quality evaluation index to obtain the deduplication result corresponding to the image set. The clustering quality evaluation index is used to reflect the cluster quality of each initial cluster.
[0169] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0170] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the image deduplication method provided by the above methods, the method comprising: determining a set of images to be deduplicated, and extracting image features of each image in the image set; clustering the image set based on the image features of each image to obtain multiple initial clusters; performing hierarchical clustering adjustment on each initial cluster based on a clustering quality evaluation index to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster.
[0171] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the image deduplication method provided by the above methods. The method includes: determining a set of images to be deduplicated and extracting image features of each image in the image set; clustering the image set based on the image features of each image to obtain multiple initial clusters; performing hierarchical clustering adjustment on each initial cluster based on a clustering quality evaluation index to obtain the deduplication result corresponding to the image set; wherein the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster.
[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image deduplication method, characterized in that, include: Determine the set of images to be deduplicated, and extract the image features of each image in the set; Based on the image features of each image, the image set is clustered to obtain multiple initial clusters; Based on the clustering quality assessment index, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set. The clustering quality assessment index is used to reflect the cluster quality of each initial cluster; The step of performing hierarchical clustering adjustment on each initial cluster based on clustering quality assessment metrics to obtain the deduplication result corresponding to the image set includes: Based on the clustering quality assessment indicators, including intra-cluster compactness, inter-cluster separation, and cluster profile coefficient, an adjustment strategy is determined. Based on the adjustment strategy, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set; The adjustment strategy includes splitting and merging; the intra-cluster compactness is characterized by the cluster diameter, and the inter-cluster separation is characterized by the inter-cluster distance; the determination of the adjustment strategy based on the intra-cluster compactness, inter-cluster separation, and cluster profile coefficient in the clustering quality assessment indicators includes: If the cluster profile coefficient of any initial cluster is less than the loose threshold and the cluster diameter of any initial cluster is greater than the diameter threshold, the adjustment strategy for any initial cluster is determined to be splitting. If the inter-cluster distance between any initial cluster and any adjacent initial cluster is less than a compactness threshold, and the difference between the cluster profile coefficient of any initial cluster and the cluster profile coefficient of any adjacent initial cluster is less than a difference threshold, then the adjustment strategy for any initial cluster and any adjacent initial cluster is determined to be merging.
2. The image deduplication method according to claim 1, characterized in that, The step of performing hierarchical clustering adjustment on each initial cluster based on the adjustment strategy to obtain the deduplication result corresponding to the image set includes: Each initial cluster with the adjustment strategy of splitting is split, and two adjacent initial clusters with the adjustment strategy of merging are merged to obtain multiple candidate clusters; The candidate clusters are updated to the initial clusters, and the updated initial clusters are adjusted hierarchically until the optimization termination condition is met. The optimization termination condition includes that the cluster profile coefficients of the initial clusters are all greater than or equal to the loose threshold and less than or equal to the compact threshold. The initial clusters that satisfy the optimization termination condition are taken as target clusters, and the deduplication result corresponding to the image set is determined based on the multiple target clusters.
3. The image deduplication method according to claim 1 or 2, characterized in that, The step of performing hierarchical clustering adjustment on each initial cluster based on clustering quality assessment metrics to obtain the deduplication result corresponding to the image set includes: Based on the clustering quality evaluation index, hierarchical clustering adjustment is performed on each initial cluster to obtain multiple target clusters; The target images in each target cluster are determined based on the distance between each image in each target cluster and the cluster center. Based on the target images in each target cluster and the cluster labels of the target images, the deduplication result corresponding to the image set is determined.
4. The image deduplication method according to claim 1 or 2, characterized in that, The extraction of image features from each image in the image set includes: Each image in the image set is input into the feature extraction model, which extracts features from each image to obtain the image features of each image output by the feature extraction model. The feature extraction model is trained based on the feature similarity between sample image features in positive samples and the similarity between sample image features in negative samples; the positive samples include different sample augmented images of the same sample image, and the negative samples include sample augmented images of different sample images.
5. The image deduplication method according to claim 1 or 2, characterized in that, Based on the image features of each image, the image set is clustered to obtain multiple initial clusters, including: Determine the initial range of cluster values; Based on the elbow algorithm, a candidate cluster value range is determined from the initial cluster value range; Based on the silhouette coefficient method, the optimal initial cluster value is determined from the range of candidate cluster values; Based on the image features of each image, the image set is clustered according to the optimal initial cluster value to obtain multiple initial clusters.
6. An image deduplication device, characterized in that, include: An extraction unit is used to determine the set of images to be deduplicated and to extract the image features of each image in the set of images; A clustering unit is used to cluster the image set based on the image features of each image to obtain multiple initial clusters; The deduplication unit is used to perform hierarchical clustering adjustment on each initial cluster based on the clustering quality evaluation index to obtain the deduplication result corresponding to the image set; the clustering quality evaluation index is used to reflect the cluster quality of each initial cluster. The step of performing hierarchical clustering adjustment on each initial cluster based on clustering quality assessment metrics to obtain the deduplication result corresponding to the image set includes: Based on the clustering quality assessment indicators, including intra-cluster compactness, inter-cluster separation, and cluster profile coefficient, an adjustment strategy is determined. Based on the adjustment strategy, hierarchical clustering adjustment is performed on each initial cluster to obtain the deduplication result corresponding to the image set; The adjustment strategy includes splitting and merging; the intra-cluster compactness is characterized by the cluster diameter, and the inter-cluster separation is characterized by the inter-cluster distance; the determination of the adjustment strategy based on the intra-cluster compactness, inter-cluster separation, and cluster profile coefficient in the clustering quality assessment indicators includes: If the cluster profile coefficient of any initial cluster is less than the loose threshold and the cluster diameter of any initial cluster is greater than the diameter threshold, the adjustment strategy for any initial cluster is determined to be splitting. If the inter-cluster distance between any initial cluster and any adjacent initial cluster is less than a compactness threshold, and the difference between the cluster profile coefficient of any initial cluster and the cluster profile coefficient of any adjacent initial cluster is less than a difference threshold, then the adjustment strategy for any initial cluster and any adjacent initial cluster is determined to be merging.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the image deduplication method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image deduplication method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Face image recognition method and device, electronic equipment and storage medium
CN109829433A