Image clustering method, electronic device and computer-readable storage medium
By using anchor maps and regular terms to build a similar matrix in multimodal image clustering and clustering based on weights, the problem of low clustering accuracy of multimodal image is solved, and a higher clustering accuracy is achieved.
Patent Information
- Application Number
- CN202510204562.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In the prior art, when clustering multimodal images, the clustering accuracy is not high due to the influence of image differences at different perspectives.
By obtaining multiple sample map collections, selecting a preset number of anchor maps, building a similarity matrix for each mode, and clustering the sample maps based on weights and regular terms to determine their categories.
The accuracy of multimodal image clustering is improved, and better clustering results are obtained by reducing the sensitivity of feature similarity to noise or outliers, and avoiding the occurrence of ordinary solutions.
Smart Images

Figure CN119693665B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular to an image clustering method, an electronic device and a computer-readable storage medium. Background Art
[0002] With the advent of the data age, a large number of images need to be processed and clustered into corresponding categories in order to manage the images. In the prior art, clustering images from a single perspective has a better clustering accuracy, that is, clustering a single modality image set is more accurate, but when clustering a multi-modality image set, the image clustering accuracy is not high due to the difference in images from different perspectives. In view of this, how to improve the accuracy of multi-modal image clustering has become an urgent problem to be solved. Summary of the invention
[0003] The main technical problem solved by the present application is to provide an image clustering method, an electronic device and a computer-readable storage medium, which can improve the accuracy of multimodal image clustering.
[0004] To solve the above technical problems, the first aspect of the present application provides an image clustering method, comprising: obtaining multiple sample image sets, and selecting a preset number of anchor images from each of the sample image sets; wherein each of the sample image sets corresponds to a modality; based on the feature similarity between the sample images in the sample image set and the anchor images, and a first regularization term corresponding to the feature similarity, constructing a similarity matrix matched by each modality; based on the similarity matrix matched by each modality, the weight of each modality and the second regularization term corresponding to the weight, clustering the sample images in all the sample image sets to determine the category corresponding to the sample images.
[0005] To solve the above technical problems, the second aspect of the present application provides an electronic device, which includes: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described in the first aspect.
[0006] In order to solve the above technical problem, the third aspect of the present application provides a computer-readable storage medium on which program data is stored. When the program data is executed by a processor, the method described in the first aspect is implemented.
[0007] The above scheme obtains the sample graph set corresponding to each modality, obtains multiple sample graph sets, and selects a preset number of sample graphs from each sample graph set as anchor graphs. The feature similarity between the sample graph and the anchor graph is obtained, and the order of magnitude of the feature similarity is effectively reduced. Based on the feature similarity between the sample graph and the anchor graph in the sample graph set and the first regularization term set for the feature similarity, the feature similarity between the sample graph and the anchor graph is updated, so that the original feature similarity between the sample graph and the anchor graph is used as a basis, the sensitivity of the similarity to noise or outliers is reduced, and the first regularization term is used to avoid the appearance of trivial solutions, so as to obtain a better similarity and construct a similarity matrix matched by each modality. Based on the similarity matrix matched by each modality, the weight of each modality and the second regularization term corresponding to the weight, the sample images in the sample image set corresponding to all modalities are clustered, so as to set the corresponding weight for each modality, so that each modality can produce its own influence in the clustering process according to the weight, avoid mutual influence between modalities, and use the second regularization term to avoid the occurrence of trivial solutions, obtain accurate clustering results, determine the categories corresponding to the sample images, and improve the accuracy of multimodal image clustering. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0009] Figure 1 It is a flowchart of an implementation method of the image clustering method of the present application;
[0010] Figure 2 It is a flowchart of another implementation method of the image clustering method of the present application;
[0011] Figure 3 It is a structural schematic diagram of an embodiment of the electronic device of the present application;
[0012] Figure 4 It is a structural schematic diagram of an implementation method of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments, and different implementation methods can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0014] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0015] The image clustering method provided in the present application is used to cluster multimodal images, that is, to cluster images collected from multiple perspectives, and its corresponding execution subject is a processing unit capable of performing image processing.
[0016] See also Figure 1 , Figure 1 : is a flow chart of an implementation method of an image clustering method of the present application, the method comprising:
[0017] S101: Acquire multiple sample image sets, and select a preset number of anchor images from each sample image set; wherein each sample image set corresponds to a modality.
[0018] Specifically, a sample graph set corresponding to each modality is obtained to obtain multiple sample graph sets, and a preset number of sample graphs are selected from each sample graph set as anchor graphs.
[0019] It should be noted that the target corresponds to sample images collected from multiple perspectives, and the sample images collected from each perspective constitute a sample image set, and each perspective corresponds to its own modality. Therefore, each sample image set corresponds to a modality.
[0020] In some implementation scenarios, multiple sample graph sets are obtained, a preset number of selected anchor graphs set for the sample graph sets is determined, and a balanced K-means and hierarchical K-means (BKHK) algorithm is used to select a preset number of sample graphs from the sample graph sets as anchor graphs matched by the corresponding sample graph sets.
[0021] In some implementation scenarios, multiple sample graph sets are obtained, a preset number of selected anchor graphs set for the sample graph sets is determined, and a preset number of sample graphs are selected from the sample graph sets using a K-means clustering algorithm as anchor graphs matched by the corresponding sample graph sets.
[0022] It can be understood that the anchor image selected in the sample image set is a representative image. For the entire sample image set, it is possible to avoid determining feature similarities between sample images. The feature similarity can be determined by comparing the sample image with the anchor image, thereby effectively reducing the order of magnitude of feature similarity.
[0023] S102: Based on the feature similarity between the sample graph and the anchor graph in the sample graph set and the first regularization term corresponding to the feature similarity, a similarity matrix matched by each modality is constructed.
[0024] Specifically, the feature similarity between the sample image and the anchor image is obtained to effectively reduce the order of magnitude of the feature similarity. Based on the feature similarity between the sample image and the anchor image in the sample image set and the first regularization term set for the feature similarity, the feature similarity between the sample image and the anchor image is updated to construct a similarity matrix matching each modality.
[0025] It should be noted that the original feature similarity between the sample image and the anchor image is used as the basis to reduce the sensitivity of the similarity to noise or outliers, and the first regular term is used to avoid the occurrence of trivial solutions, obtain better similarity, and construct a similarity matrix matching each modality.
[0026] In some implementation scenarios, the feature similarity between the sample image and the anchor image is obtained, the feature difference between the sample image and the anchor image is determined, and the conversion weight is determined based on the feature similarity and its corresponding first regularization term. The feature difference is transformed using the conversion weight to obtain the conversion deviation. The sum of the conversion deviations corresponding to the sample image and each anchor image is minimized as the matrix optimization objective. Based on the matrix optimization objective, the feature similarity between the sample image and the anchor image is updated to obtain the target similarity, and the target similarity is used to construct a similarity matrix matching the corresponding modality.
[0027] In some implementation scenarios, the feature similarity between the sample image and the anchor image is obtained, and based on the feature similarity and its corresponding first regularization term, the deviation value matching the feature similarity is indexed, and the sum of the deviation values corresponding to the sample image and each anchor image is minimized as the matrix optimization objective. Based on the matrix optimization objective, the feature similarity between the sample image and the anchor image is updated to obtain the target similarity, and the target similarity is used to construct a similarity matrix matching the corresponding modality.
[0028] Optionally, after obtaining the feature similarity between the sample images and the anchor point images in the sample image set, they are sorted in descending order according to the feature similarity, and a specified number of sample images with the highest order are obtained to construct a similarity matrix, so that a sparse solution can be obtained when the feature similarity is updated based on the optimization objective, thereby further reducing the complexity of the similarity matrix.
[0029] S103: Based on the similarity matrix matched by each modality, the weight of each modality and the second regular term corresponding to the weight, cluster the sample graphs in all sample graph sets to determine the categories corresponding to the sample graphs.
[0030] Specifically, based on the similarity matrix matched by each modality, the weight of each modality and the second regular term corresponding to the weight, the sample graphs in the sample graph set corresponding to all modalities are clustered to determine the categories corresponding to the sample graphs.
[0031] It should be noted that a corresponding weight is set for each modality so that each modality can have its own influence in the clustering process according to the weight, avoiding mutual influence between modalities, and supplemented by the second regular term to avoid the occurrence of trivial solutions, obtain accurate clustering results, determine the category corresponding to the sample image, and improve the accuracy of multimodal image clustering.
[0032] In some implementation scenarios, based on the similarity matrix corresponding to all modalities, the classification probability of the sample images belonging to different categories is determined, the reference probability corresponding to all sample images is obtained, and the classification probability of the sample images of the corresponding modalities belonging to different categories is adjusted using the weights corresponding to the corresponding modalities and their corresponding second regularization terms. The minimization of the deviation between the adjusted classification probability and the reference probability is used as the clustering optimization goal. Based on the clustering optimization goal, the classification probability, reference probability and the weights of all modalities are updated until the clustering optimization goal converges. Based on the final classification probability, the category corresponding to the sample image is determined.
[0033] In some implementation scenarios, based on the similarity matrices corresponding to all modalities, sample graphs of all modalities are clustered to obtain reference clustering results, wherein the reference clustering results include the category center of each category, and the reference graph corresponding to each category is obtained. The reference clustering results are adjusted using the weight corresponding to each modality and its corresponding second regularization term, and the deviation between the category center corresponding to the adjusted reference clustering result and the reference graph of the corresponding category is minimized as the clustering optimization target. Based on the clustering optimization target, the reference clustering results and the weights of all modalities are updated until the clustering optimization target converges, and the category corresponding to the sample graph is determined based on the final reference clustering result.
[0034] The above scheme obtains the sample graph set corresponding to each modality, obtains multiple sample graph sets, and selects a preset number of sample graphs from each sample graph set as anchor graphs. The feature similarity between the sample graph and the anchor graph is obtained, and the order of magnitude of the feature similarity is effectively reduced. Based on the feature similarity between the sample graph and the anchor graph in the sample graph set and the first regularization term set for the feature similarity, the feature similarity between the sample graph and the anchor graph is updated, so that the original feature similarity between the sample graph and the anchor graph is used as a basis, the sensitivity of the similarity to noise or outliers is reduced, and the first regularization term is used to avoid the appearance of trivial solutions, so as to obtain a better similarity and construct a similarity matrix matched by each modality. Based on the similarity matrix matched by each modality, the weight of each modality and the second regularization term corresponding to the weight, the sample images in the sample image set corresponding to all modalities are clustered, so as to set the corresponding weight for each modality, so that each modality can produce its own influence in the clustering process according to the weight, avoid mutual influence between modalities, and use the second regularization term to avoid the occurrence of trivial solutions, obtain accurate clustering results, determine the categories corresponding to the sample images, and improve the accuracy of multimodal image clustering.
[0035] See also Figure 2 , Figure 2 : is a flow chart of another embodiment of the image clustering method of the present application, the method comprising:
[0036] S201: Acquire multiple sample graph sets and determine a preset number corresponding to anchor graphs.
[0037] Specifically, a sample graph set corresponding to each modality is obtained to obtain multiple sample graph sets, and a preset number of selected anchor point graphs pre-set for the sample graph sets is determined.
[0038] S202: For each sample graph set, divide the sample graph set into a preset number of clusters based on a preset number, and obtain a sample graph matching the cluster center corresponding to each cluster as an anchor graph.
[0039] Specifically, for each sample graph set, the sample graph set is divided into a preset number of clusters according to a preset number, a cluster center is determined from each of the preset number of clusters, and the sample graph matched by the cluster center is used as the anchor graph to improve the accuracy of the anchor graph and ensure that the anchor graph is sufficiently representative.
[0040] In some implementation scenarios, for each sample graph set, the sample graph set is divided into two clusters of even size, and then each subcluster is divided into two clusters of even size, and the process is repeated layer by layer until the number of leaf nodes reaches the preset number corresponding to the anchor graph, and finally the sample graph matching the cluster center of each cluster in the last layer is used as the anchor graph.
[0041] S203: Based on the feature similarity between the sample graph and the anchor graph in the sample graph set and the first regularization term corresponding to the feature similarity, a similarity matrix matched by each modality is constructed.
[0042] Specifically, the feature similarity between the sample image and the anchor image is obtained to effectively reduce the order of magnitude of the feature similarity. Based on the feature similarity between the sample image and the anchor image in the sample image set and the first regularization term set for the feature similarity, the feature similarity between the sample image and the anchor image is updated to construct a similarity matrix matching each modality.
[0043] In some implementation scenarios, based on the feature similarity between the sample graph in the sample graph set and the anchor graph, and the first regularization term corresponding to the feature similarity, a similarity matrix matching each modality is constructed, including performing the following steps for each modality: obtaining the feature difference and feature similarity between the sample graph in the sample graph set and each anchor graph; minimizing the cumulative value corresponding to all reference sum values obtained when comparing the sample graph with each anchor graph as the matrix optimization target corresponding to the similarity matrix; wherein the reference sum value corresponds to the product between the mutually matching feature difference and feature similarity, and the sum between the first regularization term corresponding to the feature similarity; based on the matrix optimization target, the feature similarity between the sample graph and the anchor graph is updated to obtain the target similarity, and the target similarity is used to construct the similarity matrix matching the corresponding modality.
[0044] Specifically, image features are extracted from sample images in the sample image set, and the image features of the sample images in the sample image set are compared with the image features of each anchor image to determine the feature difference and feature similarity between the sample image and each anchor image.
[0045] Furthermore, the product of the mutually matching feature difference and the feature similarity is obtained, and the product is summed with the first regularization term corresponding to the feature similarity to obtain a reference sum value. The accumulated value corresponding to all reference sum values obtained when the sample graph is compared with each anchor graph is minimized as the matrix optimization target corresponding to the similarity matrix.
[0046] It can be understood that the feature similarity between the sample graph and the anchor graph is updated according to the matrix optimization objective so that the matrix optimization objective is converged, and finally the target similarity between the sample graph and the anchor graph is determined, and the target similarity is used to construct a similarity matrix matching the corresponding mode, thereby improving the accuracy of the similarity matrix.
[0047] For ease of explanation, the matrix optimization objective is expressed using the following formula:
[0048]
[0049] in, represents the total number of anchor graphs, Representation sample graph With anchor chart The feature similarity of is the regularization parameter, corresponds to the first regularization term.
[0050] It should be noted that the similarity matrix obtained by directly optimizing the solution to the above problem is non-sparse. Therefore, the sample graph set can be screened, and a sparse solution of the similarity matrix can be obtained by setting the number of sample neighbors k.
[0051] Optionally, each anchor graph is matched with a specified number of sample graphs for determining a reference and a value, and the specified number of sample graphs are obtained from a sample graph set based on the following steps, including: sorting the sample graphs in the sample graph set based on feature differences and feature similarities between the sample graphs in the sample graph set and the anchor graph, and obtaining a specified number of sample graphs with top rankings from the sorted sample graph set.
[0052] Specifically, based on the feature difference and feature similarity between the sample graph in the sample graph set and the anchor graph, the sample graphs in the sample graph set are arranged in order of feature difference from small to large, that is, feature similarity from large to small, and a specified number of sample graphs with the top order are obtained from the sorted sample graph set, so that a sparse solution can be obtained when the feature similarity is updated based on the optimization objective, thereby further reducing the complexity of the similarity matrix.
[0053] It is understandable that the optimization objective of the similarity matrix above is independent of each , so it can be split into:
[0054]
[0055] The above formula is Taking the derivative we get: ,make , obviously ,against The different values of are as follows:
[0056] First, when , , record it as ,and , The second-order derivative of is: Therefore, when When minimizing The solution is The root of .
[0057] Second, when hour, It is also minimized The optimal solution of .
[0058] In summary, the sparse solution of the similarity matrix between each modal sample graph set and a preset number of anchor graphs is:
[0059]
[0060] in, Represents K-nearest neighbor classification (KNN).
[0061] It should be noted that the first regularization term is obtained based on the information entropy of feature similarity, and the first regularization term corresponds to a regularization term parameter, the regularization term parameter is negatively correlated with the feature similarity, and the first regularization term corresponding to the regularization term parameter is multiplied to obtain the updated first regularization term for obtaining the reference and value.
[0062] Specifically, when the feature difference between the sample image and the anchor image is smaller, that is, the feature similarity between the sample image and the anchor image is higher, the regularization item parameter is smaller, so that the regularization item parameter is used to adjust the first regularization item with high precision to improve the accuracy of the target similarity.
[0063] In a specific implementation scenario, the regularization term parameter The formula is: In other specific implementation scenarios, the regularization parameter You can also set the value of linear or nonlinear change according to the negative correlation with the feature similarity.
[0064] It can be understood that the similarity matrix corresponding to each mode can be obtained by the above method .
[0065] S204: Based on the similarity matrix matched by each modality, the weight of each modality and the second regular term corresponding to the weight, cluster the sample graphs in all sample graph sets to determine the categories corresponding to the sample graphs.
[0066] Specifically, based on the similarity matrix matched by each modality, the weight of each modality and the second regular term corresponding to the weight, the sample graphs in the sample graph set corresponding to all modalities are clustered to determine the categories corresponding to the sample graphs.
[0067] In some implementation scenarios, based on the similarity matrix matched by each modality, the weight of each modality and the second regularization term corresponding to the weight, the sample graphs in all sample graph sets are clustered to determine the categories corresponding to the sample graphs, including: obtaining the classification probability matrix and the classification indication matrix matched by all modalities, and constructing the clustering objective function corresponding to each modality based on the classification probability matrix, the similarity matrix matched by each modality and the weight of each modality; wherein the classification probability matrix includes the probability that each sample graph belongs to different categories, and the classification indication matrix includes the category to which each sample graph belongs; minimizing the classification deviation and the accumulated value corresponding to the classification and value matched by each modality is used as the clustering optimization target corresponding to all modalities; wherein the classification deviation corresponds to the deviation of the classification probability matrix compared to the classification indication matrix, and the classification and value corresponds to the product of the weight of the corresponding modality and the clustering objective function, and the sum of the second regularization term corresponding to the weight; based on the clustering optimization target, the classification probability matrix, the classification indication matrix and the weight of each modality are updated, and the categories indicated by the updated classification indication matrix are used to cluster the sample graphs in all sample graph sets to determine the categories corresponding to the sample graphs.
[0068] Specifically, the classification probability matrix and classification indicator matrix of all modal matches are obtained, wherein the classification probability matrix is initialized to an arbitrary orthogonal matrix, the classification indicator matrix is initialized to an arbitrary non-negative matrix, the classification probability matrix includes the probability that each sample image belongs to a different category, the classification indicator matrix includes the category to which each sample image belongs, and each row of the classification indicator matrix has only one element that is 1, and the rest of the elements are 0.
[0069] Furthermore, based on the classification probability matrix, the similarity matrix matched by each modality and the weight of each modality, a clustering objective function corresponding to each modality is constructed, wherein the clustering objective function is used to solve the classification probability matrix.
[0070] It should be noted that based on the classification probability matrix, the similarity matrix matched by each modality and the weight of each modality, a clustering objective function corresponding to each modality is constructed, including: based on the similarity matrix matched by each modality, a Laplace matrix corresponding to each modality is constructed; based on the classification probability matrix, the Laplace matrix corresponding to each modality and the weight of each modality, a clustering objective function corresponding to each modality is constructed.
[0071] Specifically, based on the similarity matrix matched by each mode, the Laplacian matrix corresponding to each mode is constructed. The Laplacian matrix is an important matrix for describing a graph in graph theory, which reveals the connection relationship between nodes in the graph and the structural characteristics of the graph.
[0072] It should be noted that after obtaining the anchor graphs of all modes, The adjacency matrix of a mode can be defined as: ,in, is a diagonal matrix, its The diagonal elements are No. The adjacency matrix constructed by this method is is doubly random and sparse. Then, The Laplace matrix of the modes is: .
[0073] Furthermore, based on the classification probability matrix, the Laplace matrix corresponding to each mode and the weight of each mode, a clustering objective function corresponding to each mode is constructed, so as to construct a clustering objective function that can obtain the probability that the sample graph belongs to each category.
[0074] It should be noted that the clustering optimization objective corresponds to minimizing the classification deviation and the cumulative value of the classification and value of each modality match. Among them, the classification deviation corresponds to the deviation of the classification probability matrix compared to the classification indicator matrix, and the classification and value corresponds to the product of the weight of the corresponding modality and the clustering objective function, and the sum of the second regularization term corresponding to the weight.
[0075] It can be understood that the classification probability matrix, the classification indicator matrix and the weight of each mode are updated according to the clustering optimization objective so that the clustering optimization objective is converged, and the categories indicated by the updated classification indicator matrix are used to cluster the sample graphs in the set of all sample graphs to determine the categories corresponding to the sample graphs, thereby obtaining the categories corresponding to the sample graphs without the need for additional clustering analysis operations.
[0076] For ease of explanation, the clustering optimization objective is expressed as follows:
[0077]
[0078] in, , in the first item Corresponding to the clustering objective function, represents the sum of diagonal elements, represents the classification probability matrix, Indicates The second term is to simplify the clustering calculation process, introducing a discrete classification indicator matrix , can directly represent the category of the sample graph. That is Each row of has only one element that is 1, and the rest are 0. The third term is to avoid the appearance of trivial solutions, that is, to avoid the weight of each mode approaching 0. is a hyperparameter that controls the weights of different modes.
[0079] It should be noted that the classification probability matrix, the classification indicator matrix and the weight of each modality are updated based on the clustering optimization objective, including: fixing the classification indicator matrix and the weight of each modality, and updating the classification probability matrix based on the clustering optimization objective; fixing the classification probability matrix and the weight of each modality, and updating the classification indicator matrix based on the clustering optimization objective; fixing the classification probability matrix and the classification indicator matrix, and updating the weight of each modality based on the clustering optimization objective.
[0080] Specifically, the convergence process of the clustering optimization objective is divided into three steps. The first step includes: fixing the classification indicator matrix and the weight of each mode, transforming the clustering optimization objective, and updating the classification probability matrix based on the clustering optimization objective. The second step includes: fixing the classification probability matrix and the weight of each mode, transforming the clustering optimization objective, and updating the classification indicator matrix based on the clustering optimization objective. The third step includes: fixing the classification probability matrix and the classification indicator matrix, transforming the clustering optimization objective, and updating the weight of each mode based on the clustering optimization objective. The above three steps are iterated until the clustering optimization objective converges.
[0081] It is understandable that before iteration, all variables and parameters need to be initialized. Initialize to any orthogonal matrix, Initialized to any non-negative matrix, Initialize to the formula 1 / V.
[0082] First, fix and ,renew , the above clustering optimization objective can be transformed into:
[0083]
[0084] in, , the above formula can be further transformed into:
[0085]
[0086] The above formula is a standard quadratic form and can be solved based on conventional solution methods, which will not be described in detail in this application.
[0087] Secondly, fix and ,renew , the above clustering optimization objective can be transformed into:
[0088]
[0089] Among them, we can get The solution is:
[0090]
[0091] Again, fixed and ,renew , the above clustering optimization objective can be transformed into:
[0092]
[0093] in, The above formula is Taking the derivative we get:
[0094]
[0095] Let the above formula equal to 0, we can get Solution:
[0096]
[0097] It can be understood that the above steps are iterated until the clustering optimization target converges, and finally the classification indicator matrix Directly obtain the category of the sample image.
[0098] In this embodiment, for each sample graph set, the sample graph set is divided into a preset number of clusters according to a preset number, a cluster center is determined from each of the preset number of clusters, and the sample graph matched by the cluster center is used as the anchor graph, thereby improving the accuracy of the anchor graph and ensuring that the anchor graph has sufficient representativeness. The feature difference and feature similarity between the sample graph and each anchor graph are obtained, the product between the mutually matched feature difference and feature similarity is obtained, and the reference sum value is obtained by summing the product with the first regularization term corresponding to the feature similarity, and the cumulative value corresponding to all reference sum values obtained when minimizing the sample graph and each anchor graph is used as the matrix optimization target corresponding to the similarity matrix, and the feature similarity between the sample graph and the anchor graph is updated according to the matrix optimization target, so that the matrix optimization target is converged, and finally the target similarity between the sample graph and the anchor graph is determined, and the similarity matrix matched by the corresponding modality is constructed using the target similarity, thereby improving the accuracy of the similarity matrix. Based on the similarity matrix matched by each mode, the Laplace matrix corresponding to each mode is constructed. Based on the classification probability matrix, the Laplace matrix corresponding to each mode and the weight of each mode, the clustering objective function corresponding to each mode is constructed, so as to construct a clustering objective function that can obtain the probability that the sample graph belongs to each category. The clustering optimization objective corresponds to minimizing the classification deviation and the cumulative value corresponding to the classification and value matched by each mode, wherein the classification deviation corresponds to the deviation of the classification probability matrix compared to the classification indicator matrix, and the classification and value corresponds to the product of the weight of the corresponding mode and the clustering objective function, and the sum of the second regularization term corresponding to the weight. According to the clustering optimization objective, the classification probability matrix, the classification indicator matrix and the weight of each mode are updated to converge the clustering optimization objective. The categories indicated by the updated classification indicator matrix are used to cluster the sample graphs in all sample graph sets to determine the categories corresponding to the sample graphs, so that the categories corresponding to the sample graphs can be obtained without additional clustering analysis operations.
[0099] See also Figure 3 , Figure 3 It is a structural diagram of an embodiment of an electronic device of the present application, wherein the electronic device 30 includes a memory 301 and a processor 302 coupled to each other, wherein the memory 301 stores program data (not shown), and the processor 302 calls the program data to implement the method in any of the above embodiments. For descriptions of related contents, please refer to the detailed description of the above method embodiments, which will not be repeated here.
[0100] See also Figure 4 , Figure 4 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 40 stores program data 400. When the program data 400 is executed by a processor, the method in any of the above embodiments is implemented. For descriptions of related contents, please refer to the detailed description of the above method embodiments, which will not be repeated here.
[0101] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present implementation scheme.
[0102] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0104] The above description is only an implementation method of the present application, and does not limit the protection scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the protection scope of the present application.
Claims
1. An image clustering method, characterized in that: The method comprises: Acquire multiple sample image sets, and select a preset number of anchor images from each of the sample image sets; wherein each of the sample image sets corresponds to a modality; Based on the feature similarity between the sample graph in the sample graph set and the anchor graph, and the first regularization term corresponding to the feature similarity, construct a similarity matrix matched by each modality; Based on the similarity matrix matched by each modality, the weight of each modality and the second regularization term corresponding to the weight, cluster all sample graphs in the sample graph set to determine the category corresponding to the sample graph; specifically including: obtaining the classification probability matrix and classification indication matrix matched by all modalities, and constructing the clustering objective function corresponding to each modality based on the classification probability matrix, the similarity matrix matched by each modality and the weight of each modality; wherein the classification probability matrix includes the probability that each sample graph belongs to a different category, and the classification indication matrix includes the category to which each sample graph belongs; minimizing the classification deviation And the accumulated value corresponding to the classification and value matched by each modality is used as the clustering optimization target corresponding to all modalities; wherein the classification deviation corresponds to the deviation of the classification probability matrix compared to the classification indication matrix, and the classification and value corresponds to the sum of the product of the weight of the corresponding modality and the clustering objective function and the second regularization term corresponding to the weight; based on the clustering optimization target, the classification probability matrix, the classification indication matrix and the weight of each modality are updated, and the categories indicated by the updated classification indication matrix are used to cluster all the sample graphs in the sample graph set to determine the categories corresponding to the sample graphs.
2. The image clustering method according to claim 1, characterized in that: The constructing a similarity matrix matched by each modality based on the feature similarity between the sample graph in the sample graph set and the anchor graph, and the first regularization term corresponding to the feature similarity, includes performing the following steps for each modality: Obtaining feature differences and feature similarities between a sample image in the sample image set and each of the anchor point images; Minimizing the cumulative value corresponding to all reference sum values obtained when comparing the sample graph with each anchor graph is used as the matrix optimization target corresponding to the similarity matrix; wherein the reference sum value corresponds to the product between the mutually matched feature difference and the feature similarity and the sum between the first regularization term corresponding to the feature similarity; Based on the matrix optimization target, the feature similarity between the sample graph and the anchor graph is updated to obtain the target similarity, and the target similarity is used to construct a similarity matrix matched by the corresponding modality.
3. The image clustering method according to claim 2, characterized in that: Each of the anchor graphs is matched with a specified number of sample graphs for determining the reference and value, and the specified number of sample graphs are obtained from the sample graph set based on the following steps, including: Based on the feature difference and feature similarity between the sample images in the sample image set and the anchor image, the sample images in the sample image set are sorted, and the designated number of sample images with top sorting are obtained from the sorted sample image set.
4. The image clustering method according to claim 2, characterized in that: The first regularization term is obtained based on the information entropy of the feature similarity, and the first regularization term corresponds to a regularization term parameter, the regularization term parameter is negatively correlated with the feature similarity, and the first regularization term corresponding to the regularization term parameter is multiplied to obtain an updated first regularization term for obtaining the reference and value.
5. The image clustering method according to claim 1, characterized in that: The step of constructing a clustering objective function corresponding to each modality based on the classification probability matrix, the similarity matrix matched by each modality, and the weight of each modality includes: Based on the similarity matrix matched by each mode, construct a Laplace matrix corresponding to each mode; Based on the classification probability matrix, the Laplace matrix corresponding to each modality and the weight of each modality, a clustering objective function corresponding to each modality is constructed.
6. The image clustering method according to claim 1, characterized in that: The updating of the classification probability matrix, the classification indication matrix and the weight of each modality based on the clustering optimization objective includes: The classification indicator matrix and the weight of each modality are fixed, and the classification probability matrix is updated based on the clustering optimization objective; Fixing the classification probability matrix and the weight of each modality, and updating the classification indication matrix based on the clustering optimization objective; The classification probability matrix and the classification indication matrix are fixed, and the weight of each modality is updated based on the clustering optimization objective.
7. The image clustering method according to claim 1, characterized in that: The step of obtaining a plurality of sample graph sets and selecting a preset number of anchor graphs from each of the sample graph sets includes: Acquire multiple sample graph sets and determine the preset number corresponding to the anchor graph; For each of the sample graph sets, the sample graph set is divided into a preset number of clusters based on the preset number, and a sample graph matched by a cluster center corresponding to each cluster is obtained as the anchor graph.
8. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Fast multi-view discrete clustering method and system based on anchor point diagram
CN114399653A