Image clustering method, electronic equipment and storage medium
By performing feature extraction and connection probability calculation on the image, screening the nearest neighbor nodes, and eliminating the negative connections of noise, the problem of negative connections in image clustering is solved, and the accuracy and reliability of clustering are improved.
Patent Information
- Application Number
- CN202510278092.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, due to the unrosity between feature nodes during image clustering, there are high similarity negative connections in the K-nearest neighbor graph, which damages the clustering effect of the graph convolutional neural network, resulting in a decrease in clustering accuracy, and these wrong connections are difficult to eliminate.
An image clustering method is proposed. By extracting the object image feature, determining the neighboring nodes of the node, and calculating the connection probability between the node and the neighboring node, filtering out the neighboring nodes with high connection probability, eliminating the negative connections of noise, and retaining the positive connections to improve the accuracy of clustering.
By optimizing the K nearest neighbor graph, eliminating the negative noise connections and retaining the positive connections, the accuracy of image clustering is significantly improved and the reliability of clustering results is ensured.
Smart Images

Figure CN120198696A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image clustering, and in particular to an image clustering method, an electronic device, and a storage medium. Background Art
[0002] In the existing technology, object images are usually subjected to feature extraction to obtain feature nodes, and then the K-nearest neighbor graph of the feature nodes is obtained and input into a graph convolutional neural network for clustering to obtain a clustering result. Due to the non-robustness between feature nodes, there are some feature nodes with high similarity negative connections in the K-nearest neighbor graph (belonging to two objects, but with high similarity, generally occurring in the case of poor picture quality). These connections will damage the clustering effect of the graph convolutional neural network. Specifically, these negative connections may act as bridges to aggregate all feature nodes belonging to two objects together, seriously affecting the clustering accuracy, and it is very difficult to eliminate such large-scale errors later. Summary of the Invention
[0003] This application provides at least an image detection method, a training method for related models, related devices, and equipment, which can improve the clustering accuracy.
[0004] In a first aspect of this application, an image clustering method is provided. The method includes: respectively performing feature extraction on a plurality of object images to obtain the original features corresponding to each object image, where each object image includes an object; respectively taking each object image as a node, and based on the original features of each node, determining the first number of nearest neighbor nodes corresponding to each node, where the feature similarity between the node and the corresponding nearest neighbor nodes meets the nearest neighbor requirement; for each node, obtaining the connection probability between the node and each corresponding nearest neighbor node, where the connection probability represents whether the node and the nearest neighbor node belong to the same object; based on the connection probability corresponding to the node, screening the first number of nearest neighbor nodes corresponding to the node to obtain the second number of nearest neighbor nodes; using the second number of nearest neighbor nodes corresponding to each node to perform clustering on the plurality of object images to obtain a clustering result, where the clustering result represents whether each object image belongs to the same object.
[0005] Among them, for each node, obtaining the connection probability between the node and each corresponding nearest neighbor node includes: respectively taking each node as the current node, and each nearest neighbor node of the current node as the current nearest neighbor node; encoding based on the original features of the current node and each current nearest neighbor node to obtain the encoded features of the current node and each current nearest neighbor node; splicing the encoded feature of the current node with the encoded features of each current nearest neighbor node to obtain the splicing features corresponding to each current nearest neighbor node; inputting the splicing features corresponding to each current nearest neighbor node into a probability prediction network for processing to obtain the connection probability between each current nearest neighbor node and the current node.
[0006] Among them, encoding is performed based on the original features of the current node and each current neighbor node to obtain the encoded features of the current node and each current neighbor node, including: stacking the original features of the current node and the original features of each current neighbor node to obtain an original feature matrix; constructing an adjacency matrix using the original feature matrix; inputting the original feature matrix and the adjacency matrix into a Transformer encoder for encoding processing to obtain an encoded feature matrix, where the encoded feature matrix includes the encoded features of the current node and each current neighbor node.
[0007] Among them, the probability prediction network is a multi-layer perceptron network; and / or, splicing the encoded features of the current node with the encoded features of each current neighbor node respectively to obtain the spliced features corresponding to each current neighbor node, including: splicing the encoded features of the current node to the end of each row of features in the encoded feature matrix respectively to obtain a spliced feature matrix, where each row of features in the encoded feature matrix is the encoded features of the current node and each current neighbor node respectively, and each row of features in the spliced feature matrix is the spliced features of the current node and each current neighbor node respectively; and, inputting the spliced features corresponding to each current neighbor node into the probability prediction network for processing to obtain the connection probability between each current neighbor node and the current node, including: inputting the spliced feature matrix into the probability prediction network for processing to obtain the connection probability between each current neighbor node and the current node.
[0008] Among them, based on the connection probability corresponding to the node, screening the first number of neighbor nodes corresponding to the node to obtain the second number of neighbor nodes, including: taking each node as the current node respectively, and each neighbor node of the current node as the current neighbor node; inputting the connection probability between each current neighbor node and the current node into the threshold prediction network for processing to obtain the probability threshold corresponding to the current node; screening the nodes with connection probabilities greater than the probability threshold from the first number of current neighbor nodes to obtain the second number of current neighbor nodes.
[0009] Among them, screening the first number of neighbor nodes corresponding to the node based on the connection probability corresponding to the node to obtain the second number of neighbor nodes further includes: calculating the density values of the current node and each current neighbor node based on the connection probability; screening the nodes with connection probabilities greater than the probability threshold from the first number of current neighbor nodes to obtain the second number of current neighbor nodes, including: selecting the second number of current neighbor nodes from the first number of current neighbor nodes whose density values are higher than the density value of the current node and whose connection probabilities are greater than the probability threshold.
[0010] Among them, the threshold prediction network is an LSTM network; and / or, the density value of the node is the sum of the connection probabilities between the node and its corresponding neighbor nodes.
[0011] Among them, the method further includes: obtaining a plurality of sample images, where each sample image includes a sample object, some sample images contain the same sample object, and some sample images contain different sample objects; respectively performing feature extraction on each sample image to obtain the sample original features corresponding to each sample image, where each sample image serves as a sample node; based on the sample original features of each sample node, determining the first number of sample neighbor nodes corresponding to each sample node, where the feature similarity between the sample node and the corresponding sample neighbor nodes meets the neighbor requirement; for each sample node, obtaining the sample connection probability between the sample node and each corresponding sample neighbor node, where the sample connection probability represents whether the sample node and the sample neighbor node belong to the same sample object; inputting the sample connection probability between the sample node and each corresponding sample neighbor node into a threshold prediction network for processing to obtain the predicted probability threshold corresponding to the sample node; and adjusting the threshold prediction network based on the difference between the predicted probability threshold corresponding to the sample node and the labeled probability threshold.
[0012] Among them, before inputting the sample connection probability between the sample node and each corresponding sample neighbor node into a threshold prediction network for processing to obtain the sample probability threshold corresponding to the sample node, the method further includes: calculating the sample density values of the sample node and each corresponding sample neighbor node based on the sample connection probability; selecting, from the first number of sample neighbor nodes corresponding to the sample node, the sample neighbor nodes whose sample density values are higher than the sample density value of the sample node to obtain the third number of sample neighbor nodes; respectively assigning different values to the predicted probability threshold of the sample node, and for different assignments, selecting, from the third number of sample neighbor nodes, the fourth number of sample neighbor nodes whose sample connection probabilities are higher than the assignment; using the fourth number, the third number, and the first number corresponding to each assignment to calculate the effect index value corresponding to each assignment; and based on the effect index values corresponding to each assignment, determining the optimal assignment from each assignment, and taking the optimal assignment as the labeled probability threshold corresponding to the sample connection probability of the sample node and the corresponding first number of neighbor nodes.
[0013] Among them, using the fourth number, the third number, and the first number corresponding to each assignment to calculate the effect index value corresponding to each assignment includes: calculating the accuracy rate and recall rate corresponding to each assignment using the fourth number, the third number, and the first number corresponding to each assignment; and calculating the effect index value corresponding to each assignment based on the accuracy rate and recall rate corresponding to each assignment; and / or, based on the effect index values corresponding to each assignment, determining the optimal assignment from each assignment includes: performing linear fitting on the effect index values corresponding to each assignment to obtain a curve of the effect index changing with the probability threshold, and selecting the probability threshold corresponding to the optimal effect index value from the curve as the optimal assignment.
[0014] The second aspect of the present application provides an electronic device, including a memory and a processor coupled to each other. The processor is configured to execute program instructions stored in the memory to implement the image clustering method in the first aspect above.
[0015] The third aspect of the present application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the image clustering method in the first aspect above is implemented.
[0016] In the above solution, by separately extracting features from multiple object images to obtain the original features corresponding to each object image, taking each object image as a node, and according to the original features of each node, determining the first number of neighbor nodes whose node feature similarity meets the neighbor requirement for each node, then calculating the connection probability between the node and each corresponding neighbor node respectively, and then using the connection probability corresponding to the node to screen the first number of neighbor nodes corresponding to the node to obtain the second number of neighbor nodes. Through the second number of neighbor nodes corresponding to each node, multiple object images are clustered to obtain a clustering result to determine whether each object image belongs to the same object. Before clustering the first number of neighbor nodes corresponding to each object image, the first number of neighbor nodes are purified using the connection probability corresponding to the image, and the connection probability represents whether the node and the neighbor node belong to the same object. In this way, some negative connection neighbor points with high feature similarity can be eliminated, and at the same time, positive connection neighbor points with low feature similarity can also be retained, so as to significantly improve the accuracy of clustering in the subsequent clustering process.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.
[0019] Figure 1 is a schematic flowchart of an embodiment of the image clustering method of the present application;
[0020] Figure 2 is a schematic flowchart of an embodiment of the training method of the threshold prediction network of the present application;
[0021] Figure 3 is a schematic flowchart of another embodiment of the image clustering method of the present application;
[0022] Figure 4 is a schematic framework diagram of an embodiment of the image clustering device of the present application;
[0023] Figure 5It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;
[0024] Figure 6 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0025] The following will combine the accompanying drawings of the specification to elaborate on the solutions of the embodiments of the present application in detail.
[0026] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0027] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, the term "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple types or any combination of at least two of multiple types. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0028] In the fields of image processing and computer vision, image clustering is a crucial task. However, in practical applications, due to the influence of complex scenarios such as illumination changes, pose changes, occlusions, etc., it leads to the instability of image feature extraction and the inaccuracy of clustering results. Especially between low-quality object images, there is a situation of high-similarity negative connections, making the traditional cosine similarity measurement method ineffective. Therefore, this proposal presents an image clustering algorithm based on a Transformer encoder, aiming to improve the robustness and accuracy of object clustering.
[0029] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the image clustering method of the present application. Specifically, it may include the following steps:
[0030] Step S110: Extract features from multiple object images respectively to obtain the original features corresponding to each object image, where each object image includes an object.
[0031] The present application is mainly applied to the field of image clustering. After extracting features from object images and determining the K-nearest neighbor graph of the nodes corresponding to the object images, the K-nearest neighbor graph of the nodes is purified to eliminate the noise negative pair connections in the K-nearest neighbor graph and retain sufficient hard example positive pair connections, ensuring the recall rate and accuracy in subsequent graph neural network clustering.
[0032] The object images in this application can be acquired by an image acquisition device, such as those acquired by devices like cameras, cameras, mobile phones, etc., or can also be images stored in a database. The objects contained in the object images can be cars, pets, etc., and no specific limitations are made here.
[0033] In this application, multiple object images are of the same category. If they are of different categories, they can be visually distinguished. For example, an image of a dog and an image of a car can be easily distinguished. Therefore, multiple object images in this application are of the same category. For example, they are all images of vehicles, or they are all images of cats and dogs.
[0034] After obtaining multiple object images, for all object images, use a feature extraction network (such as ResNet, VGG, etc.) to extract the original features. In addition, before extracting the original features, use an object detection algorithm (such as MTCNN, FaceBoxes, etc.) to obtain the specific positions of the objects in the object images, and align and crop the detected objects to eliminate the influence of pose changes on feature extraction. Assume the number of objects is N, and the feature of node i corresponding to the object image is fi, and the feature dimension of fi is D.
[0035] For example, when the object images all include vehicles, determine the positions of the vehicles in the object images, and align and crop the detected vehicles to retain the images of the front or rear parts of the vehicles. Usually, there are vehicle logos on both the front and rear of the vehicles, and the vehicle models are also marked on the rear. Therefore, using the images of the front or rear parts of the vehicles for feature extraction can reduce the workload of feature extraction and improve efficiency.
[0036] In addition, when the object image contains a person, the face position in the object image can be aligned and cropped.
[0037] Step S120: Respectively take each object image as a node, and based on the original features of each node, determine the first number of nearest neighbor nodes corresponding to each node, where the feature similarity between the node and the corresponding nearest neighbor nodes meets the nearest neighbor requirements.
[0038] In some embodiments, an object image can be taken as a node, and the original feature corresponding to the node is an element in the node, such as node f i =(A i1 , …, A iD ) where A iD represents the original feature. Select a node as the current node, calculate the feature similarity between the current node and other nodes, and take the other nodes that meet the nearest neighbor requirements as the first number of nearest neighbor nodes of this node. For example, feature similarity calculation methods such as cosine similarity and Euclidean distance can be used.
[0039] Specifically, the cosine similarity is used to calculate the feature similarity between the current node and other nodes, and the obtained feature similarities are sorted from large to small. The first number of other nodes corresponding to the feature similarities in the sorting are selected as the neighbor nodes of the current node.
[0040] Step S130: For each node, obtain the connection probability between the node and its corresponding neighbor nodes, where the connection probability indicates whether the node and the neighbor node belong to the same object.
[0041] In some embodiments, a neighbor threshold can be set to determine whether the neighbor node and the node belong to the same object through the neighbor threshold. Specifically, the neighbor threshold can be set to 0.8. If the feature similarity of the first number of neighbor nodes of the current node is greater than 0.8, then the feature similarity of the neighbor node is used as the connection probability of the neighbor node; otherwise, the connection probability of the neighbor node is set to 0.
[0042] In other embodiments, it can be assumed that there is a certain probability distribution relationship between the connection probability between the node and its corresponding neighbor nodes and their feature similarities. For example, machine learning algorithms such as logistic regression and support vector machines can be used to fit this relationship, so as to predict the connection probability according to the feature similarity.
[0043] In other embodiments, in complex scenarios, it is easy to have the situation of high similarity negative connections between low-quality object images. If calculated using feature similarity, it will seriously interfere with the clustering effect. Therefore, a new way is needed to measure the similarity between nodes. Considering that the Transfomer encoder was first applied to the field of natural language processing, and its core lies in the self-attention layer, which can learn the complex relationships between various words in a sentence. Therefore, in this embodiment, the Transfomer encoder can be used to mine the neighbor mutual relationships of each node and predict the probability that two nodes belong to the same object. Specifically, refer to steps S131 to S134.
[0044] Step S131: Respectively take each node as the current node, and the neighbor nodes of the current node are the current neighbor nodes.
[0045] Step S132: Encode based on the original features of the current node and each current neighbor node to obtain the encoded features of the current node and each current neighbor node.
[0046] In some embodiments, the original features of the current node and the original features of each current neighbor node are stacked to obtain the original feature matrix F i ∈R (K+1)×D K represents the first number. Then use the original feature matrix Fi , construct an adjacency matrix such as the adjacency matrix The calculation method can be Then, input the original feature matrix F i and the adjacency matrix into the Transformer encoder for encoding processing together to obtain the encoded feature matrix F' i , where the encoded feature matrix F' i includes the encoded features of the current node and each current neighboring node. The Transformer encoder in this embodiment is composed of multi-head attention (Multi-Head Attention), and its number of layers can be adjusted according to actual needs.
[0047] Step S133: Concatenate the encoded feature of the current node with the encoded features of each current neighboring node respectively to obtain the concatenated features corresponding to each current neighboring node.
[0048] In some embodiments, the encoded feature of the current node can be concatenated to the end of each row of features in the encoded feature matrix respectively to obtain a concatenated feature matrix, where each row of features in the encoded feature matrix is the encoded feature of the current node and each current neighboring node respectively, and each row of features in the concatenated feature matrix is the concatenated feature of the current node and each current neighboring node respectively.
[0049] Specifically, the 0th row feature of the encoded feature matrix F' i is the encoded feature of the current node i after encoding. Concatenate this encoded feature to the end of each row of features of the encoded feature matrix F' i to obtain a feature matrix with dimension R (K+1)×2D . The purpose of concatenation is to obtain the encoded feature of the current node, so that it has the opportunity to perform matrix multiplication with the encoded features of other neighboring nodes, similar to the cosine similarity calculation method, so as to obtain the concatenated features corresponding to each current neighboring node.
[0050] Step S134: Input the concatenated features corresponding to each current neighboring node into the probability prediction network for processing to obtain the connection probability between each current neighboring node and the current node.
[0051] In some embodiments, the concatenated feature matrix can be input into the probability prediction network for processing to obtain the connection probability c ij between each current neighboring node and the current node. When the probability prediction network is a multi-layer perceptron network (MultilayerPerceptron, MLP), the concatenated feature matrix can be input into the MLP for processing to obtain the connection probability c ij between each current neighboring node and the current node.
[0052] In addition, the probability prediction network can also be a Bayesian network or the like, which is not specifically limited herein.
[0053] Step S140: Based on the connection probabilities corresponding to the nodes, screen the first number of neighboring nodes corresponding to the nodes to obtain the second number of neighboring nodes.
[0054] In this step, the connection probabilities in the previous step are used to optimize the first number of neighboring nodes of the nodes, eliminate the noise negative pair connections, and retain sufficient positive pair connections to ensure the recall rate and accuracy of subsequent clustering.
[0055] In some embodiments, the threshold prediction network can be used to optimize the first number of neighboring nodes of the nodes. Specifically, refer to Step S141 to Step S143.
[0056] Step S141: Respectively take each node as the current node, and each neighboring node of the current node as the current neighboring node.
[0057] Step S142: Input the connection probabilities between each current neighboring node and the current node into the threshold prediction network for processing to obtain the probability threshold corresponding to the current node.
[0058] In some embodiments, the threshold prediction network is a trained network. After inputting the connection probabilities between the current node and its corresponding neighboring nodes into the threshold prediction network, the probability threshold corresponding to the current node will be output. Among them, the threshold prediction network can be an LSTM network.
[0059] Step S143: Select the nodes with connection probabilities greater than the probability threshold from the first number of current neighboring nodes to obtain the second number of current neighboring nodes.
[0060] In some embodiments, based on the connection probabilities, the density values of the current node and each current neighboring node can be calculated. For example, the density value of a node is the sum of the connection probabilities between the node and its corresponding neighboring nodes: d i is the density value. Select the second number of current neighboring nodes from the first number of current neighboring nodes whose density values are higher than the density value of the current node and whose connection probabilities are greater than the probability threshold. Thus, the optimization of the first number of neighboring nodes of the nodes is completed.
[0061] Step S150: Use the second number of neighboring nodes corresponding to each node to cluster multiple object images to obtain a clustering result, where the clustering result indicates whether each object image belongs to the same object.
[0062] In some embodiments, input the second number of neighboring nodes corresponding to each node into a trained graph neural network for clustering to obtain a clustering result.
[0063] Input the second number of purified neighboring nodes into the trained graph neural network for clustering, thereby improving the accuracy of clustering.
[0064] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the training method of the threshold prediction network of the present application. Specifically, it may include the following steps:
[0065] Step S210: Obtain a plurality of sample images, where each sample image includes a sample object, some sample images contain the same sample object, and some sample images contain different sample objects.
[0066] Regarding the sample images, they can be obtained from a sample database, and the ratio between the positive sample images and the negative sample images in the sample images is preset.
[0067] Step S220: Extract features from each sample image respectively to obtain the sample original features corresponding to each sample image, where each sample image serves as a sample node.
[0068] After obtaining the sample images, use a feature extraction network to extract features from them to obtain the sample original features corresponding to each sample image.
[0069] Step S230: Based on the sample original features of each sample node, determine the first number of sample neighboring nodes corresponding to each sample node, where the feature similarity between the sample node and the corresponding sample neighboring node meets the neighboring requirement.
[0070] In some embodiments, the cosine similarity can be used to calculate the feature similarity between a sample node and other sample nodes, and the obtained feature similarities are sorted, and the other sample nodes corresponding to the first number of the largest feature similarities are selected as sample neighboring nodes.
[0071] Step S240: For each sample node, obtain the sample connection probability between the sample node and the corresponding sample neighboring nodes, where the sample connection probability represents whether the sample node and the sample neighboring node belong to the same sample object.
[0072] In some embodiments, each sample node is respectively used as the current sample node, and each sample neighbor node of the current sample node is the current sample neighbor node. The sample original features of the current sample node and the sample original features of each current sample neighbor node are stacked to obtain a sample original feature matrix. A sample adjacency matrix is constructed using the sample original feature matrix. Then, the sample original feature matrix and the sample adjacency matrix are input into a Transformer encoder for encoding processing to obtain a sample encoded feature matrix. The sample encoded features of the current sample node are respectively concatenated to the end of each row of features in the sample encoded feature matrix to obtain a sample concatenated feature matrix. Then, the sample concatenated feature matrix is input into a probability prediction network for processing to obtain the sample connection probabilities between the sample node and its corresponding sample neighbor nodes.
[0073] Step S250: Input the sample connection probabilities between the sample node and its corresponding sample neighbor nodes into a threshold prediction network for processing to obtain the predicted probability threshold corresponding to the sample node.
[0074] In some embodiments, the sample connection probabilities between the sample node and its corresponding sample neighbor nodes can be input into a threshold prediction network for processing to obtain the predicted probability threshold corresponding to the sample node. Among them, the threshold prediction network can be an LSTM network.
[0075] Step S260: Adjust the threshold prediction network based on the difference between the predicted probability threshold corresponding to the sample node and the labeled probability threshold.
[0076] In some embodiments, the difference between the predicted probability threshold corresponding to the sample node and the labeled probability threshold can be calculated using a loss function to obtain a loss value, and the parameters of the threshold prediction network are adjusted through the loss value. After multiple iterations, the iteration stops until the threshold prediction network meets the expectation.
[0077] In other embodiments, the labeled probability threshold can be obtained through the effect index value. Specifically, please refer to Step S261 to Step S265.
[0078] Step S261: Based on the sample connection probabilities, calculate the sample density values of the sample node and its corresponding sample neighbor nodes.
[0079] In some embodiments, the sum of the sample connection probabilities of the sample node and its corresponding sample neighbor nodes can be used as the sample density value.
[0080] Step S262: Select the sample neighbor nodes whose sample density values are higher than the sample density value of the sample node from the first number of sample neighbor nodes corresponding to the sample node to obtain the third number of sample neighbor nodes.
[0081] Step S263: Different values are assigned to the prediction probability thresholds of the sample nodes respectively. For different assignments, from the third number of sample neighbor nodes, the fourth number of sample neighbor nodes with sample connection probabilities higher than the assignment value are selected.
[0082] Step S264: Using the fourth number, the third number, and the first number corresponding to each assignment, the effectiveness index values corresponding to each assignment are calculated.
[0083] In some embodiments, using the fourth number, the third number, and the first number corresponding to each assignment, the accuracy rate and the recall rate corresponding to each assignment are calculated. Then, based on the accuracy rate and the recall rate corresponding to each assignment, the effectiveness index values corresponding to each assignment are calculated.
[0084] Specifically, the effectiveness index can be F Score , where P is the accuracy rate and R is the recall rate. In this embodiment, starting from the effectiveness index, the index where, P t , R t respectively represent the accuracy rate and the recall rate of the sample neighbor nodes whose sample density values are higher than the sample density value of the sample node and whose similarity is higher than the prediction probability threshold t after assignment, 0 < t < 1. Therefore, the prediction probability threshold t of the sample node can be assigned a value of 0.3, the first number is 100, and there are 80 sample neighbor nodes that are the same sample object as the sample node. Using the assignment of 0.3 to screen the sample connection probabilities of these 100 sample neighbor nodes, the fourth number is obtained as 50, and among the 50, 47 sample neighbor nodes are the same sample object as the sample node, and 3 are not the same sample object as the sample node. Therefore, using 47 / 100 = 0.47, it is the accuracy rate P t ; 3 / 50 = 0.06, which is the recall rate R t . Then, different assignments can be made to the prediction probability threshold t. For example, the assignments are: 0.1, 0.2, 0.4, 0.51, 0.52. After multiple assignments, multiple corresponding effectiveness index values F Score can be obtained.
[0085] Step S265: Based on the effectiveness index values corresponding to each assignment, the optimal assignment is determined from each assignment, and the optimal assignment is used as the annotation probability threshold corresponding to the sample connection probability of the sample node and the corresponding first number of neighbor nodes.
[0086] In some embodiments, linear fitting is performed on the effect index values corresponding to each assignment to obtain a curve showing how the effect index changes with the probability threshold, that is, E(t) is a curve. The probability threshold corresponding to the optimal effect index value is selected from the curve as the optimal assignment, and the optimal assignment is used as the annotation probability threshold corresponding to the sample connection probability of the sample node and the corresponding first number of neighboring nodes. The loss value is calculated using this annotation probability threshold and the predicted probability threshold, thereby adjusting the network parameters of the threshold prediction network.
[0087] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of another embodiment of the image clustering method of the present application.
[0088] Specifically, the following steps may be included:
[0089] Step S310: Feature extraction is respectively performed on multiple object images to obtain the original features corresponding to each object image.
[0090] This step is the same as step 110 above and will not be elaborated here.
[0091] Step S320: Each object image is respectively used as a node, and based on the original features, the connection probability between the node and other nodes is calculated.
[0092] In some embodiments, each object image is used as a node, and the Markov chain is used to calculate the transition probability when the node transfers to other nodes, and the connection probability between two nodes is obtained through multiple iterations.
[0093] Step S330: Other nodes whose connection probabilities meet the preset connection requirements are selected as the first number of neighboring nodes of the node.
[0094] In some embodiments, a connection probability threshold can be set to screen neighboring nodes. For example, other nodes with a connection probability greater than the connection probability threshold can be used as the first number of neighboring nodes of the node.
[0095] Step S340: Based on the connection probability corresponding to the node, the first number of neighboring nodes corresponding to the node are screened to obtain the second number of neighboring nodes.
[0096] This step is the same as step 140 above and will not be elaborated here.
[0097] Step S350: Using the second number of neighboring nodes corresponding to each node, multiple object images are clustered to obtain a clustering result.
[0098] This application optimizes the K-nearest neighbor graph of the object images, eliminates the noise negative pair connections, and retains sufficient positive pair connections, ensuring the recall rate and accuracy of subsequent graph neural network clustering. To achieve this advantage, starting from the final metric F_Score of the clustering algorithm, the optimal assignment is solved. Such a problem can be transformed into a problem of finding the maximum value of F_Score for time series data. By training the LSTM network, the optimal threshold can be obtained. And the transfomer encoder is used to mine the neighbor mutual relationship of each neighbor node, predict the probability that two nodes belong to the same class, and as a more robust similarity measure, replace the cosine similarity.
[0099] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and constitutes any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0100] Please refer to Figure 4 , Figure 4 which is a schematic framework diagram of an embodiment of the image clustering device 400 of this application. The image clustering device 400 includes: a feature extraction module 410, a screening module 420, a probability calculation module 430, an optimization module 440, and a clustering module 450. The feature extraction module 410 performs feature extraction on multiple object images respectively to obtain the original features corresponding to each object image, where each object image includes an object. The screening module 420 performs taking each object image as a node respectively, and based on the original features of each node, determines the first number of neighbor nodes corresponding to each node, where the feature similarity between the node and the corresponding neighbor node meets the neighbor requirement. The probability calculation module 430 performs for each node, obtaining the connection probability between the node and each corresponding neighbor node, where the connection probability represents whether the node and the neighbor node belong to the same object. The optimization module 440 performs screening the first number of neighbor nodes corresponding to the node based on the connection probability corresponding to the node to obtain the second number of neighbor nodes. The clustering module 450 performs clustering the multiple object images by using the second number of neighbor nodes corresponding to each node to obtain a clustering result, where the clustering result represents whether each object image belongs to the same object.
[0101] In some embodiments, the probability calculation module 430 performs, for each node, obtaining the connection probability between the node and each corresponding neighboring node, including: respectively taking each node as the current node, and each neighboring node of the current node as the current neighboring node; encoding based on the original features of the current node and each current neighboring node to obtain the encoded features of the current node and each current neighboring node; concatenating the encoded features of the current node with the encoded features of each current neighboring node respectively to obtain the concatenated features corresponding to each current neighboring node; and inputting the concatenated features corresponding to each current neighboring node into a probability prediction network for processing to obtain the connection probability between each current neighboring node and the current node.
[0102] In some embodiments, the probability calculation module 430 performs encoding based on the original features of the current node and each current neighboring node to obtain the encoded features of the current node and each current neighboring node, including: stacking the original features of the current node and the original features of each current neighboring node to obtain an original feature matrix; constructing an adjacency matrix using the original feature matrix; and inputting the original feature matrix and the adjacency matrix into a transformer encoder for encoding processing to obtain an encoded feature matrix, where the encoded feature matrix includes the encoded features of the current node and each current neighboring node.
[0103] In some embodiments, the probability prediction network is a multi-layer perceptron network; and / or, the probability calculation module 430 performs concatenating the encoded features of the current node with the encoded features of each current neighboring node respectively to obtain the concatenated features corresponding to each current neighboring node, including: concatenating the encoded features of the current node to the end of each row of features in the encoded feature matrix respectively to obtain a concatenated feature matrix, where each row of features in the encoded feature matrix is the encoded feature of the current node and each current neighboring node, and each row of features in the concatenated feature matrix is the concatenated feature of the current node and each current neighboring node; and inputting the concatenated features corresponding to each current neighboring node into a probability prediction network for processing to obtain the connection probability between each current neighboring node and the current node, including: inputting the concatenated feature matrix into a probability prediction network for processing to obtain the connection probability between each current neighboring node and the current node.
[0104] In some embodiments, the optimization module 440 performs screening the first number of neighboring nodes corresponding to a node based on the connection probability corresponding to the node to obtain the second number of neighboring nodes, including: respectively taking each node as the current node, and each neighboring node of the current node as the current neighboring node; inputting the connection probability between each current neighboring node and the current node into a threshold prediction network for processing to obtain the probability threshold corresponding to the current node; and screening the nodes with a connection probability greater than the probability threshold from the first number of current neighboring nodes to obtain the second number of current neighboring nodes.
[0105] In some embodiments, the optimization module 440 performs screening on the first number of neighboring nodes corresponding to a node based on the connection probability corresponding to the node to obtain the second number of neighboring nodes, further including: calculating the density values of the current node and each current neighboring node based on the connection probability; screening, from the first number of current neighboring nodes, the nodes with a connection probability greater than the probability threshold to obtain the second number of current neighboring nodes, including: selecting, from the first number of current neighboring nodes, the second number of current neighboring nodes whose density value is higher than that of the current node and whose connection probability is greater than the probability threshold.
[0106] In some embodiments, the threshold prediction network is an LSTM network; and / or, the density value of the node executed by the optimization module 440 is the sum of the connection probabilities between the node and its corresponding neighboring nodes.
[0107] In some embodiments, the optimization module 440 further performs: obtaining a plurality of sample images, where each sample image includes a sample object, some sample images contain the same sample object, and some sample images contain different sample objects; respectively performing feature extraction on each sample image to obtain the sample original features corresponding to each sample image, where each sample image serves as a sample node; determining the first number of sample neighboring nodes corresponding to each sample node based on the sample original features of each sample node, where the feature similarity between the sample node and the corresponding sample neighboring node meets the neighboring requirement; for each sample node, obtaining the sample connection probability between the sample node and its corresponding sample neighboring nodes, where the sample connection probability represents whether the sample node and the sample neighboring node belong to the same sample object; inputting the sample connection probability between the sample node and its corresponding sample neighboring nodes into the threshold prediction network for processing to obtain the predicted probability threshold corresponding to the sample node; adjusting the threshold prediction network based on the difference between the predicted probability threshold corresponding to the sample node and the labeled probability threshold.
[0108] In some embodiments, before the optimization module 440 inputs the sample connection probability between the sample node and each corresponding sample neighbor node into the threshold prediction network for processing to obtain the sample probability threshold corresponding to the sample node, the method further includes: calculating, based on the sample connection probability, the sample density values of the sample node and each corresponding sample neighbor node; selecting, from the first number of sample neighbor nodes corresponding to the sample node, the sample neighbor nodes whose sample density values are higher than the sample density value of the sample node to obtain the third number of sample neighbor nodes; respectively assigning different values to the predicted probability threshold of the sample node, and for different assignments, selecting, from the third number of sample neighbor nodes, the fourth number of sample neighbor nodes whose sample connection probabilities are higher than the assignment; using the fourth number, the third number, and the first number corresponding to each assignment to calculate the effect index value corresponding to each assignment; and based on the effect index values corresponding to each assignment, determining the optimal assignment from each assignment, and using the optimal assignment as the labeled probability threshold corresponding to the sample connection probability of the sample node and the first number of neighbor nodes corresponding thereto.
[0109] In some embodiments, the optimization module 440 executes calculating the effect index value corresponding to each assignment by using the fourth number, the third number, and the first number corresponding to each assignment, including: calculating the accuracy rate and recall rate corresponding to each assignment by using the fourth number, the third number, and the first number corresponding to each assignment; calculating the effect index value corresponding to each assignment based on the accuracy rate and recall rate corresponding to each assignment; and / or, based on the effect index values corresponding to each assignment, determining the optimal assignment from each assignment, including: performing linear fitting on the effect index values corresponding to each assignment to obtain a curve of the effect index varying with the probability threshold, and selecting the probability threshold corresponding to the optimal effect index value from the curve as the optimal assignment.
[0110] Please refer to Figure 5 , Figure 5 is a schematic framework diagram of an embodiment of the electronic device 50 of the present application. The electronic device 50 includes a memory 51 and a processor 52 that are coupled to each other. The processor 52 is configured to execute program instructions stored in the memory 51 to implement the steps in any of the above-described embodiments of the image clustering method. In a specific implementation scenario, the electronic device 50 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 50 may further include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0111] Specifically, the processor 52 is used to control itself and the memory 51 to implement the steps in any of the above image clustering method embodiments. The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip with the ability to process signals. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 52 may be implemented jointly by integrated circuit chips.
[0112] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the computer-readable storage medium 60 of the present application. The computer-readable storage medium 60 stores program instructions 601 that can be run by a processor, and the program instructions 601 are used to implement the steps in any of the above image clustering method embodiments.
[0113] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0114] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0115] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0116] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically as individual units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
Claims
1. An image clustering method, characterized in that: include: Extracting features from a plurality of object images respectively to obtain original features corresponding to each of the object images, wherein each of the object images includes an object; Taking each of the object images as a node, determining a first number of neighboring nodes corresponding to each of the nodes based on original features of each of the nodes, wherein the feature similarity between the node and the corresponding neighboring node meets the neighboring requirement; For each of the nodes, obtaining a connection probability between the node and each of the corresponding neighboring nodes, wherein the connection probability indicates whether the node and the neighboring node belong to the same object; Based on the connection probability corresponding to the node, a first number of neighboring nodes corresponding to the node are screened to obtain a second number of neighboring nodes; The plurality of object images are clustered using a second number of neighboring nodes corresponding to each of the nodes to obtain a clustering result, wherein the clustering result indicates whether each of the object images belongs to the same object.
2. The method according to claim 1, characterized in that The step of obtaining, for each of the nodes, a connection probability between the node and each of the corresponding neighboring nodes includes: Each of the nodes is taken as a current node, and each of the neighboring nodes of the current node is taken as a current neighboring node; Encoding is performed based on original features of the current node and each of the current neighboring nodes to obtain encoding features of the current node and each of the current neighboring nodes; Concatenate the encoding features of the current node with the encoding features of each of the current neighboring nodes to obtain concatenated features corresponding to each of the current neighboring nodes; The splicing features corresponding to each of the current neighboring nodes are input into a probability prediction network for processing to obtain the connection probability between each of the current neighboring nodes and the current node.
3. The method according to claim 2, characterized in that The encoding based on the original features of the current node and each of the current neighboring nodes to obtain the encoding features of the current node and each of the current neighboring nodes includes: Stacking the original features of the current node and the original features of each of the current neighboring nodes to obtain an original feature matrix; Using the original feature matrix, construct an adjacency matrix; The original feature matrix and the adjacency matrix are input into a transformer encoder for encoding processing to obtain an encoding feature matrix, wherein the encoding feature matrix includes encoding features of the current node and each of the current neighboring nodes.
4. The method according to claim 2, characterized in that: The probability prediction network is a multi-layer perceptron network; And / or, the coding features of the current node are spliced with the coding features of each of the current neighboring nodes to obtain the spliced features corresponding to each of the current neighboring nodes, including: splicing the coding features of the current node to the end of each row of features in the coding feature matrix to obtain a spliced feature matrix, wherein each row of features in the coding feature matrix is the coding features of the current node and each of the current neighboring nodes, and each row of features in the spliced feature matrix is the spliced features of the current node and each of the current neighboring nodes; and, inputting the spliced features corresponding to each of the current neighboring nodes into the probability prediction network for processing to obtain the connection probability between each of the current neighboring nodes and the current node, including: inputting the spliced feature matrix into the probability prediction network for processing to obtain the connection probability between each of the current neighboring nodes and the current node.
5. The method according to claim 1, characterized in that The step of screening a first number of neighboring nodes corresponding to the node based on the connection probability corresponding to the node to obtain a second number of neighboring nodes includes: Each of the nodes is taken as a current node, and each of the neighboring nodes of the current node is taken as a current neighboring node; Inputting the connection probability between each of the current neighboring nodes and the current node into a threshold prediction network for processing to obtain a probability threshold corresponding to the current node; From the first number of current neighbor nodes, nodes whose connection probability is greater than the probability threshold are screened to obtain a second number of current neighbor nodes.
6. The method according to claim 5, characterized in that The method further comprises screening a first number of neighboring nodes corresponding to the node based on the connection probability corresponding to the node to obtain a second number of neighboring nodes: Based on the connection probability, density values of the current node and each of the current neighboring nodes are calculated; The step of selecting nodes whose connection probability is greater than the probability threshold from the first number of current neighbor nodes to obtain a second number of current neighbor nodes includes: From the first number of current neighbor nodes, a second number of current neighbor nodes whose density values are higher than the density value of the current node and whose connection probabilities are greater than the probability threshold are selected.
7. The method according to claim 6, characterized in that The threshold prediction network is an LSTM network; And / or, the density value of the node is the sum of the connection probabilities between the node and corresponding neighboring nodes.
8. The method according to claim 6, characterized in that The method further comprises: Acquire a plurality of sample images, wherein each of the sample images includes a sample object, some of the sample images include the same sample object, and some of the sample images include different sample objects; Extracting features from each of the sample images respectively to obtain original features of the samples corresponding to each of the sample images, wherein each of the sample images is used as a sample node; Based on the original sample features of each of the sample nodes, determine a first number of sample neighbor nodes corresponding to each of the sample nodes, wherein the feature similarity between the sample node and the corresponding sample neighbor node meets the neighbor requirement; For each of the sample nodes, obtaining a sample connection probability between the sample node and each of the corresponding sample neighboring nodes, wherein the sample connection probability indicates whether the sample node and the sample neighboring node belong to the same sample object; The sample connection probability between the sample node and the corresponding sample neighboring nodes is input into the threshold prediction network for processing, and the prediction probability threshold value corresponding to the sample node is obtained. The threshold prediction network is adjusted based on the difference between the prediction probability threshold and the annotation probability threshold corresponding to the sample node.
9. The method according to claim 8, characterized in that Before inputting the sample connection probability between the sample node and the corresponding sample neighboring nodes into the threshold prediction network for processing to obtain the sample probability threshold corresponding to the sample node, the method further includes: Based on the sample connection probability, the sample density value of the sample node and the corresponding sample neighboring nodes is calculated; Selecting the sample neighboring nodes whose sample density values are higher than the sample density value of the sample node from the first number of sample neighboring nodes corresponding to the sample node, to obtain a third number of sample neighboring nodes; Assigning different values to the prediction probability thresholds of the sample nodes respectively, and selecting, from the third number of sample neighbor nodes, a fourth number of sample neighbor nodes whose sample connection probabilities are higher than the assigned values according to the different assigned values; Using the fourth quantity, the third quantity and the first quantity corresponding to each of the assignments, calculate the effect index value corresponding to each of the assignments; Based on the effect index values corresponding to the assignments, an optimal assignment is determined from the assignments, and the optimal assignment is used as a labeling probability threshold corresponding to the sample connection probability of the sample node and the corresponding first number of neighboring nodes.
10. The method according to claim 9, characterized in that The calculating the effect index value corresponding to each of the assignments by using the fourth quantity, the third quantity, and the first quantity corresponding to each of the assignments includes: Calculate the precision and recall corresponding to each of the assignments using the fourth quantity, the third quantity, and the first quantity corresponding to each of the assignments; Based on the accuracy and recall rate corresponding to each of the assignments, calculate the effect index value corresponding to each of the assignments; And / or, determining the optimal assignment from the assignments based on the effect indicator values corresponding to the assignments, comprises: Linear fitting is performed on the effect index values corresponding to the assignments to obtain a curve showing the effect index changing with the probability threshold, and the probability threshold corresponding to the optimal effect index value is selected from the curve as the optimal assignment.
11. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the processor is used to execute program instructions stored in the memory to implement the image clustering method according to any one of claims 1 to 10.
12. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the image clustering method according to any one of claims 1 to 10 is implemented.