Clustering Center Determination Method, Device, Equipment and Computer Storage Medium
By combining feature learning and clustering processes in the image clustering model, the feature representation vector is optimized, which solves the problem of insufficient bucket similarity caused by poor feature learning in the prior art, and improves the accuracy of image and video retrieval.
Patent Information
- Application Number
- CN202110298057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-03-19
AI Technical Summary
In the prior art, clustering is performed based on the learned image representation when clustering, resulting in poor feature learning, resulting in insufficient similarity between buckets and affecting the accuracy of subsequent retrieval.
Multiple rounds of clustering are used to combine feature learning and clustering processes, and through feature coding, clustering loss value and model adjustment, the accuracy and clustering effect of feature representation vectors are optimized, and dynamic loss adjustment is designed to ensure the stability of embedding.
It improves the accuracy of feature learning and clustering effect, ensures the accuracy of clustering centers, and improves the accuracy of image retrieval and video retrieval.
Smart Images

Figure CN113704528B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to the field of Artificial Intelligence (AI) technology, and provides a method, apparatus, device and computer storage medium for determining cluster centers. Background Art
[0002] Large-scale image retrieval often relies on bucket retrieval. Bucket retrieval mainly divides a large amount of original data into multiple non-overlapping data subsets, and each data subset belongs to a bucket. During retrieval, only the matching samples need to be searched in the bucket that best matches the target sample, so as to improve the retrieval efficiency. Generally speaking, buckets are generated by clustering. That is, for one million samples, if they are divided into 10,000 buckets, then the cluster centers are 10,000. The effect of bucketing has a great impact on the final retrieval result. Ideally, it is expected that samples with similar features can be assigned to the same bucket, so that the recall of a certain bucket is similar to the true samples.
[0003] However, currently during clustering, clustering is usually performed based on the learned image representations to obtain cluster centers. Therefore, the accuracy of the image representations directly determines the quality of the clustering effect. When the feature learning is poor, it is very easy to cause insufficient bucket similarity, which in turn affects the subsequent retrieval accuracy. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, device and computer storage medium for determining cluster centers, which are used to improve the accuracy of image representations and the image clustering effect.
[0005] On the one hand, a method for determining cluster centers is provided. The method includes:
[0006] Obtain an image sample set, and based on the image sample set, perform multiple rounds of clustering using an image clustering model until a preset convergence condition is met. One round of clustering includes the following operations:
[0007] Perform feature encoding on each image sample in the image sample set, obtain corresponding multiple feature representation vectors based on the encoding results, and determine a feature representation loss value;
[0008] Cluster the multiple feature representation vectors, and determine a clustering loss value based on the clustering result;
[0009] When it is determined that the image clustering model has not converged based on the feature representation loss value and the clustering loss value, adjust the parameters of the image clustering model;
[0010] Output multiple cluster center vectors included in the clustering result obtained in the last round of clustering. One cluster center vector corresponds to one image category.
[0011] On the one hand, a clustering center determination device is provided. The device includes:
[0012] A sample acquisition unit for acquiring a set of image samples;
[0013] A clustering unit for performing multiple rounds of clustering on the basis of the set of image samples by using an image clustering model until a preset convergence condition is satisfied. Wherein, the clustering unit includes a feature encoding subunit, a clustering subunit, and a model adjustment subunit:
[0014] The feature encoding subunit is configured to respectively perform feature encoding on each image sample in the set of image samples, obtain corresponding multiple feature representation vectors based on the encoding results, and determine a feature representation loss value;
[0015] The clustering subunit is configured to cluster the multiple feature representation vectors and determine a clustering loss value based on the clustering results;
[0016] The model adjustment subunit is configured to adjust the parameters of the image clustering model when it is determined that the image clustering model has not converged based on the feature representation loss value and the clustering loss value;
[0017] An output unit for outputting multiple clustering center vectors included in the clustering result obtained in the last round of clustering, where one clustering center vector corresponds to one image category.
[0018] Optionally, the feature encoding subunit is specifically configured to:
[0019] For each of the image samples, perform at least one basic feature extraction respectively to obtain corresponding basic feature vectors;
[0020] Respectively perform feature compression on each of the obtained basic feature vectors to obtain the feature representation vectors corresponding to each of the image samples.
[0021] Optionally, the sample acquisition unit is further configured to:
[0022] Construct multiple groups of image samples based on the set of image samples, each group of image samples includes at least two image samples, and includes the labeled similarity between the at least two image samples;
[0023] Then the feature encoding subunit is specifically configured to:
[0024] For each of the multiple groups of image samples, perform the following operations: Based on the feature representation vectors of at least two image samples included in one group of the multiple groups of image samples, obtain the predicted similarity between the at least two image samples, and compare the obtained predicted similarity with the corresponding labeled similarity to obtain the corresponding comparison result;
[0025] Determine the feature representation loss value according to each comparison result obtained for each group of image samples.
[0026] Optionally, each group of image samples includes a first image sample, a second image sample, and a third image sample, and the similarity between the first image sample and the second image sample in each group of image samples is not less than a preset similarity threshold, and the similarity between the first image sample and the third image sample is less than the similarity threshold;
[0027] Then the feature encoding subunit is specifically configured to:
[0028] For each of the multiple groups of image samples, perform the following operations: In one group of the at least one group of image samples, obtain the first similarity between the first image sample and the second image sample, and obtain the second similarity between the first image sample and the third image sample, and based on the first similarity and the second similarity, obtain the triplet loss value corresponding to the one group of image samples;
[0029] Obtain the feature representation loss value based on the triplet loss values respectively corresponding to each of the obtained groups of image samples.
[0030] Optionally, the clustering subunit is specifically configured to:
[0031] Select at least one feature representation vector corresponding to the first image sample and the third image sample in the multiple groups of image samples from the multiple feature representation vectors;
[0032] Perform clustering on the at least one feature representation vector.
[0033] Optionally, the clustering subunit is specifically configured to:
[0034] For each of the multiple feature representation vectors, perform the following operations:
[0035] Map one feature representation vector in the multiple feature representation vectors to each image category respectively to obtain the corresponding image category vectors; where each element in the image category vector corresponds to an image category, and the value of each element represents whether the image sample corresponding to the one feature representation vector belongs to the corresponding image category;
[0036] Compare the image category vector with the image category vector determined for the image sample corresponding to the one feature representation vector in the previous round of clustering to obtain a comparison result;
[0037] Determine the clustering loss value based on the multiple comparison results obtained for the multiple feature representation vectors.
[0038] Optionally, the clustering subunit is specifically configured to:
[0039] Select at least one feature representation vector corresponding to the first image sample and the third image sample in the multiple image sample groups from the multiple feature representation vectors;
[0040] Determining the clustering loss value based on the multiple comparison results obtained for the multiple feature representation vectors includes:
[0041] Determine the clustering loss value based on the comparison results corresponding to the at least one feature representation vector.
[0042] Optionally, the clustering subunit is specifically configured to:
[0043] Determine the similarity between the one feature representation vector and each clustering center vector determined in this round of clustering respectively;
[0044] Based on the obtained similarities, determine whether the image sample corresponding to the one feature representation vector belongs to the image category represented by the corresponding clustering center vector respectively, and obtain a determination result;
[0045] Obtain the image category vector based on the determination result.
[0046] Optionally, the model adjustment subunit is further configured to:
[0047] Determine the weight value of the clustering loss value based on the difference between the feature representation loss value and the feature representation loss value of the previous round of clustering; wherein, the weight value is negatively correlated with the difference;
[0048] Determine the total model loss value of the image clustering model based on the feature representation loss value, the clustering loss value, and the weight value;
[0049] Determine whether the image clustering model converges based on the total model loss value.
[0050] Optionally, the model adjustment subunit is further configured to:
[0051] Perform scale transformation on the feature representation loss value and the clustering loss value so that the measurement scales of the feature representation loss value and the clustering loss value are the same.
[0052] Optionally, the feature encoding subunit is further configured to:
[0053] When it is determined that the feature encoding submodel has not converged based on the feature representation loss value, adjust the parameters of the feature encoding submodel;
[0054] Perform at least one iterative update on the adjusted feature encoding submodel until the current round of clustering meets the set conditions; wherein, in each iterative update, perform feature encoding on each image sample respectively based on the feature encoding submodel, determine the feature representation loss value based on the multiple feature representation vectors obtained by encoding, and adjust the parameters of the feature encoding submodel based on the feature representation loss value, and the set conditions include that the number of times of feature encoding in the current round of clustering reaches the set number threshold, or the feature encoding submodel meets the convergence condition;
[0055] Then, the clustering subunit is specifically configured to:
[0056] Perform clustering based on the multiple feature representation vectors obtained by encoding using the feature encoding submodel after the last iterative update in the current round of clustering.
[0057] Optionally, the apparatus further includes a video batching unit and a video retrieval unit.
[0058] The video batching unit is configured to perform the following operations respectively for each video in the video database: based on the multiple clustering center vectors, respectively determine the image categories to which each image frame included in one video belongs, and associate the one video with each image category according to the determined image categories.
[0059] The video retrieval unit is configured to receive a video retrieval request sent by a target user, where the video retrieval request carries a target video; based on the multiple clustering center vectors, determine the image categories to which each image frame included in the target video belongs, and determine the videos associated with the determined image categories as candidate matching videos; select a target matching video to be recommended to the target user from the candidate matching videos.
[0060] On the one hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of any of the above methods are implemented.
[0061] On the one hand, a computer storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of any of the above methods are implemented.
[0062] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any of the above methods.
[0063] In the clustering center determination method according to the embodiments of the present application, an end-to-end image clustering model is constructed. When using the image clustering model to cluster an image sample set, feature learning and the clustering process are combined together, that is, unsupervised clustering is performed while feature learning is carried out. Thus, feature learning and clustering promote and optimize each other. In other words, during the feature learning process, to a certain extent, the distribution characteristics of the previous clustering can be learned, and furthermore, the obtained feature representation vectors will also tend to certain distribution characteristics, improving the accuracy of feature learning. In addition, based on this, continuing to cluster can maintain a greater similarity between the new clustering and the previous clustering to a greater extent, making the features clustered into the same category closer to each other. Correspondingly, the finally obtained clustering center is more accurate, and thus the clustering effect is better. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0065] Figure 1 It is a schematic diagram of an application scenario provided for the embodiments of the present application;
[0066] Figure 2 It is a schematic flowchart of the clustering center determination method provided for the embodiments of the present application;
[0067] Figure 3 It is a model structure diagram of the image clustering model provided for the embodiments of the present application;
[0068] Figure 4 It is a schematic flowchart of a clustering provided for the embodiments of the present application;
[0069] Figure 5 It is a schematic structure diagram of the feature encoding sub-model provided for the embodiments of the present application;
[0070] Figure 6 It is a schematic diagram of the image category vector provided for the embodiments of the present application;
[0071] Figure 7Another schematic diagram of the clustering process provided by the embodiments of the present application;
[0072] Figure 8 A schematic diagram of the application process taking the video recommendation scenario as an example provided by the embodiments of the present application;
[0073] Figure 9 A schematic diagram of video retrieval provided by the embodiments of the present application;
[0074] Figure 10 A schematic diagram of a structure of a clustering center determination device provided by the embodiments of the present application;
[0075] Figure 11 A schematic diagram of a structure of a computer device provided by the embodiments of the present application. Detailed implementation manners
[0076] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the scope of protection of the present application. Without conflict, the embodiments in the present application and the features in the embodiments may be arbitrarily combined with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0077] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application are explained here first:
[0078] Image category recognition: Recognition is performed by considering only the category of the image (such as people, dogs, cats, birds, etc.) without considering the specific instance of the image, and the category to which the image belongs is given. A typical example is the recognition task in the large-scale general object recognition open-source dataset ImageNet, which recognizes which of the 1000 categories a certain object belongs to.
[0079] ImageNet pre-trained model: A deep learning network model is trained based on the large-scale general object recognition open-source dataset ImageNet, and the parameter weights of the model are the ImageNet pre-trained model.
[0080] OpenImage pre-trained model: A deep learning network model is trained based on the open-source dataset OpenImage, and the parameter weights of the model are the OpenImage pre-trained model.
[0081] Bucket Retrieval: The original large amount of data is first divided into multiple non-overlapping data subsets, and each data subset belongs to a bucket. When retrieving, only need to find the matching samples in the bucket that best matches the target sample. Therefore, bucket retrieval can improve the retrieval efficiency. Bucket retrieval can be applied to image retrieval. Specifically, all images in the image database can be bucket-divided according to categories. One bucket corresponds to one category. Then, when performing image retrieval, the image to be retrieved is roughly matched with each bucket, and then fine-matched in the bucket with the best matching degree, thereby improving the image recognition speed. In addition, bucket retrieval can also be applied to video bucket retrieval. Specifically, first, for each image frame of all videos in the database, find the nearest clustering center as the bucket to which the image belongs. In this way, multiple buckets to which each video belongs are obtained. Then, when retrieving a certain video, for each frame image of the video, find the nearest clustering center as the bucket to which it belongs, and use the database videos in these buckets as roughly matched videos. Then, perform further fine-matching among the roughly matched videos, and the video retrieval speed is faster.
[0082] Image Clustering: The process of dividing a set of physical or abstract objects into multiple classes composed of similar objects is called clustering. The clustering in the embodiments of this application mainly refers to image clustering, which divides a complete image set into multiple image subsets, that is, the above-mentioned various buckets. The image subsets generated by clustering are sets of multiple images, and these images are similar to each other within the same image subset and different from the images in other image subsets. Generally speaking, clustering methods can include the K-Means algorithm, the K-Medoids algorithm, partitioning methods (such as the CLARANS algorithm), and density-based clustering methods (DBSCAN), etc.
[0083] Clustering Center Vector: A special image sample in image clustering analysis, which can be used to represent a certain image category, and other image samples determine whether they belong to this image category by calculating the distance from it.
[0084] Triplet Loss: Triplet loss is a loss function in deep learning. In tasks where the training goal is to obtain the feature representation vector (embedding) of samples, triplet loss is also often used, such as the embedding of images. The input of triplet loss is a triplet, and a triplet includes an anchor (a) example, a positive (p) example, and a negative (n) example. By optimizing the distance between the anchor example and the positive example to be less than the distance between the anchor example and the negative example, the similarity calculation between samples is realized.
[0085] In image feature learning, a triplet consists of three image samples, namely image sample a, image sample p, and image sample n. Image sample a and image sample p are positive samples, that is, image sample a and image sample p are similar image samples or belong to the same category of image samples. Image sample a and image sample n are negative samples, that is, image sample a and image sample p are image samples with a relatively low similarity or belong to different categories of image samples. The ultimate optimization goal is to narrow the distance between the embeddings of image sample a and image sample p, that is, to make the first similarity between image sample a and image sample p higher, and to widen the distance between the embeddings of image sample a and image sample n, that is, to make the second similarity between image sample a and image sample n lower, that is, to make the distance difference between positive samples and negative samples as large as possible. The triplet loss value of a triplet is used to characterize the difference degree between the first similarity and the second similarity. Therefore, the objective function of the triplet loss value is defined as follows:
[0086] L = max(d(a, p) - d(a, n) + margin, 0)
[0087] Where L is the triplet loss value, d(a, p) represents the distance between image sample a and image sample p, d(a, n) represents the distance between image sample a and image sample n, and margin is a constant greater than 0.
[0088] Currently, when performing image clustering, it is usually based on the learned image representation for clustering to obtain the clustering center. Therefore, the accuracy of the image representation directly determines the quality of the clustering effect. When the feature learning is poor, it is very easy to cause insufficient bucket similarity, which in turn affects the subsequent retrieval accuracy.
[0089] Considering that the root cause of the above problems in the related technology is that the feature learning stage and the clustering stage in the related technology are relatively independent, that is, first feature learning and then clustering, so there is no mutual influence between the two stages. The independent way of the two stages severs the gradient transfer between clustering and feature learning, which in turn causes the problem that insufficient clustering similarity caused by poor feature learning affects subsequent retrieval. At the same time, the feature learning stage does not specifically learn the distribution of global samples, which easily causes the global sample features not to follow the similarity distribution. For example, the similarity of samples belonging to the same category is relatively low. Therefore, to solve the above problems, the feature learning process can be combined with the clustering process, so that the feature learning can be integrated into the clustering result, so that the learned features follow the similarity distribution to a certain extent and are more accurate, and thus more accurate clustering can be performed when clustering.
[0090] In view of this, an embodiment of the present application provides a method for determining a clustering center. In this method, an end-to-end image clustering model is constructed. When using the image clustering model to cluster an image sample set, feature learning and the clustering process are combined, that is, unsupervised clustering is performed while feature learning is carried out, so that feature learning and clustering promote and optimize each other. In other words, during the feature learning process, to a certain extent, the distribution characteristics of the previous clustering can be learned, and thus the learned feature representation vectors will also tend to certain distribution characteristics, improving the accuracy of feature learning. In addition, continuing to cluster on this basis can maintain the similarity between the new clustering and the previous clustering to a greater extent, making the features clustered into the same category closer to each other. Correspondingly, the finally obtained clustering center will be more accurate, and thus the clustering effect will be better.
[0091] At the same time, since the adjustment based on the clustering result may cause a large deviation in the result of feature learning, an embodiment of the present application designs dynamic loss adjustment to ensure the stability of the embedding, that is, the clustering weight is adjusted according to the embedding learning situation, and finally the learned embedding has a better effect in similarity clustering.
[0092] After introducing the design concept of the embodiment of the present application, the following briefly introduces the application scenarios applicable to the technical solution of the embodiment of the present application. It should be noted that the following introduced application scenarios are only used to illustrate the embodiment of the present application rather than to limit it. In the specific implementation process, the technical solution provided by the embodiment of the present application can be flexibly applied according to actual needs.
[0093] The solution provided by the embodiment of the present application can be applied to most scenarios that require image clustering, such as image retrieval or video scenarios. As Figure 1 shown, it is a schematic diagram of an application scenario provided by the embodiment of the present application. In this scenario, it may include a server 10 and a terminal 20.
[0094] Among them, the server 10 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, but is not limited thereto.
[0095] Server 10 may include one or more processors 101, a memory 102, an I / O interface 103 for interacting with other devices, etc. In addition, the server 10 may also be configured with a database 104, which can be used to store data such as model data, received target videos, video binning data, or image binning data involved in the solutions provided in the embodiments of the present application. Among them, program instructions of the clustering center determination method provided in the embodiments of the present application may be stored in the memory 102 of the server 10. When these program instructions are executed by the processor 101, they can be used to implement the steps of the clustering center determination method provided in the embodiments of the present application to obtain an image clustering center vector and video binning retrieval, etc.
[0096] The terminal 20 is any terminal device capable of providing an input and search function interface, such as a mobile phone, a tablet computer (PAD), a desktop computer, a laptop computer, a smart TV, or a wearable smart device, etc. In the terminal 20, an application that can provide an input and search function interface can be installed, such as an image retrieval application or a video application. Correspondingly, the server 10 can be an image retrieval server or a video application server. The application involved in the embodiments of the present application can be a software client, or a browser, a small program, etc. The server 10 is the background server corresponding to the software or the small program, etc., and the specific type of the client is not limited. It should be noted that when the application is a browser, the input and search functions can be specifically implemented through the web page opened by the browser. Then, the server 10 is the background server corresponding to the web page.
[0097] When the server 10 is an image retrieval server, the server 10 can use the clustering center determination method provided in the embodiments of the present application to perform feature learning on the images stored in the database 104 to obtain corresponding feature representation vectors, and at the same time cluster each image, divide all images into corresponding bins, and obtain the clustering center vector corresponding to each bin. Furthermore, when the user performs image retrieval, the user can input the image to be retrieved in the retrieval interface of the terminal 20, and then send a retrieval request to the server 10. Correspondingly, the server 10 can receive the image to be retrieved input by the user, and match the image to be retrieved with the clustering center vectors corresponding to each bin to determine the bin with the highest matching degree, and then perform further retrieval of the images in the bin to return to the user the images precisely matched in the bin.
[0098] When the server 10 is a video application server, the server 10 can divide each video in the database 104 into each bin. Specifically, each video can be represented by the image frames it contains. Furthermore, while performing feature learning on the image frames of all videos using the clustering center determination method provided in the embodiments of the present application, cluster the image frames, divide each image frame into the corresponding bin, and obtain the clustering center vector corresponding to each bin. Then, divide each video into the bin to which the image frames it includes belong. Furthermore, when a user performs a video search, the user can input the video to be searched in the search interface of the terminal 20, and then send a search request to the server 10. Correspondingly, the server 10 can receive the video to be searched input by the user, and respectively match each image frame included in the video to be searched with the clustering center vectors corresponding to each bin to determine the bin with the highest matching degree for each image frame. Then, recall the videos in these bins to obtain the recalled videos, and further perform an exact match among the recalled videos to return the videos obtained by the exact match to the user.
[0099] A direct or indirect communication connection can be established between the server 10 and the terminal 20 through one or more networks 30. The network 30 can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiments of the present application do not limit this.
[0100] Certainly, the method provided in the embodiments of the present application is not limited to Figure 1 the application scenarios shown, and can also be used in other possible application scenarios, which are not limited by the embodiments of the present application. For Figure 1 the functions that can be realized by each device in the application scenarios shown will be described together in the subsequent method embodiments, and will not be elaborated here too much. Next, the technologies involved in the embodiments of the present application will be briefly introduced.
[0101] Artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0102] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0103] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision research involves related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0104] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0105] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common applications include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0106] The solution provided in the embodiments of this application involves technologies such as computer vision and machine learning in artificial intelligence. Specifically, image processing is performed through computer vision technology, and the image clustering model is trained through machine learning technology. Specific details will be described in the following embodiments.
[0107] Please refer to Figure 2 , which is a schematic flowchart of the clustering center determination method provided by the embodiments of the present application. This method can be executed by the server 10 or the terminal 20 in Figure 1 , or can be jointly executed by the terminal 20 and the server 10. In the embodiments of the present application, specifically, taking the case where this method is executed by the server 10 as an example, the process of this method is introduced as follows.
[0108] Step 201: Obtain an image sample set.
[0109] In the embodiments of the present application, the image sample set includes multiple image samples. For example, in an image retrieval scenario, the image sample set may include images in a certain image database, or, in a video retrieval scenario, the image sample set may include image frames included in each video in a video database.
[0110] Step 202: Based on the image sample set, perform multiple rounds of clustering using an image clustering model until a preset convergence condition is met.
[0111] Refer to Figure 3 shown, which is a model structure diagram of the image clustering model provided by the embodiments of the present application. In the image clustering model, it may specifically include a feature encoding sub-model and a clustering sub-model. Since the specific execution process of this model will be introduced in detail in the subsequent content, it will not be elaborated too much here.
[0112] In the embodiments of the present application, one round of clustering may include the following operations:
[0113] Perform feature encoding on multiple image samples extracted from the image sample set respectively, obtain corresponding multiple feature representation vectors based on the encoding results, and determine the feature representation loss value; and, perform clustering on the multiple feature representation vectors, and determine the clustering loss value based on the clustering results; furthermore, determine whether the image clustering model converges based on the feature representation loss value and the clustering loss value. When it is determined that the image clustering model does not converge, adjust the parameters of the image clustering model to enter the next round of clustering process.
[0114] Step 203: Output multiple clustering center vectors included in the clustering result obtained in the last round of clustering, and one clustering center vector corresponds to one image category.
[0115] In the embodiments of the present application, when the image clustering model meets the preset convergence condition, the clustering is completed, that is, each image sample in the image sample set is divided into each category. Correspondingly, the feature representation vectors that can represent each category in the clustering result can be used as the clustering center vectors for each category, and thus multiple clustering center vectors can be obtained. One clustering center vector corresponds to one image category, and this clustering center vector can be used in subsequent image retrieval or video retrieval applications.
[0116] Next, the specific process of clustering by the image clustering model will be introduced.
[0117] In the embodiments of the present application, the process of clustering by the image clustering model can essentially be understood as the process of training the image clustering model. When the image clustering model completes clustering, the image clustering model meets the convergence condition. During the training process of the image clustering model, multiple training processes are required to make the image clustering model gradually tend to converge, and each training process is similar to one round of clustering process. Therefore, the following mainly takes one round of clustering as an example for introduction.
[0118] See Figure 4 shown in the figure, which is a schematic flowchart of clustering by the image clustering model.
[0119] Step 401: Construct multiple image sample groups based on the image sample set. Each image sample group includes at least two image samples and the labeled similarity between at least two image samples.
[0120] In the embodiments of the present application, the constructed image sample groups are mainly used in the feature learning process. Different image sample groups can be constructed based on different feature encoding submodels. For example, when the feature encoding submodel is a model based on supervised prediction, each constructed image sample group can include at least two image samples, and each image sample group includes at least two image samples and the labeled similarity between at least two image samples.
[0121] For example, the feature encoding submodel can also be a model based on sample similarity prediction. Then each image sample group can include two image samples, and label annotation is performed for each image sample group. The label indicates that the two image samples included in each image sample group are similar samples or dissimilar samples, and then supervised feature learning is performed based on the label.
[0122] Alternatively, when the feature encoding sub-model is also a model based on triplet loss, each constructed group of image samples is a triplet, which is a triplet of labeled data for embedding training. The triplet is divided into (a, p, n), representing the anchor image, positive image, and negative image respectively. Among them, a and p are similar images or images of the same category, and a and n are dissimilar images or images of different categories. Such triplets need to be labeled before training the image clustering model.
[0123] In the embodiments of the present application, the image samples in the group of image samples are used for both embedding learning and clustering. For example, when the group of image samples is a triplet, in each round of clustering, since a and p among the three samples a, p, and n belong to the same category and n belongs to a different category, if a and p also participate in the calculation of the clustering loss, then the sampling quantity of the category to which a and p belong is twice that of n. This will cause the head category to have an unbalanced learning of the category where n is due to too many samples. Therefore, a and n in each triplet can be selected to participate in the calculation of the clustering loss, or p and n can be selected to participate in the calculation of the clustering loss to balance the positive and negative sample clustering losses and avoid unbalanced category learning caused by too many samples in the head category.
[0124] Step 402: Perform feature encoding on each image sample respectively, and obtain corresponding multiple feature representation vectors based on the encoding results.
[0125] In the embodiments of the present application, refer to Figure 3 As shown, in each round of clustering, input the multiple groups of image samples participating in this round of clustering into the image clustering model. Among them, the image samples included in the multiple groups of image samples participating in this round of clustering can include all image samples in the image sample set, or can be selected from the image sample set. If selection from the image sample set is required, it can be randomly selected from the image sample set based on the set selection quantity, or can be selected from the image sample set according to the set selection rules.
[0126] Specifically, the feature encoding sub-model used for feature encoding can be any model that can encode to obtain feature representation vectors. For example, a convolutional neural network (Convolutional Neural Networks, CNN) or a residual network (Residual Network, ResNet) can be used, such as a feature encoding sub-model based on basic embedding models such as resnet101, resnet34, Densenet, mobilenet, or GoogleNet. Of course, other possible model structures can also be used, and the embodiments of the present application do not limit this.
[0127] See Figure 5 As shown, it is a schematic structural diagram of a feature encoding sub - model provided by an embodiment of the present application. Among them, the feature encoding sub - model can include a basic feature extraction layer and a feature compression layer. The basic feature extraction layer performs at least one basic feature extraction on each image sample respectively to obtain corresponding basic feature vectors. Then, the feature compression layer respectively performs feature compression on each obtained basic feature vector to obtain a feature representation vector corresponding to each image sample.
[0128] In the embodiment of the present application, the basic feature extraction layer can adopt any model capable of performing image basic feature extraction, such as resnet101.
[0129] In practical applications, in order to improve the convergence speed of the image clustering model, a pre - trained basic feature extraction model can be adopted, that is, the basic feature parameters of the pre - trained basic feature extraction model can be applied to the image clustering model provided by the embodiment of the present application. See Figure 5 As shown, it is a schematic structure shown with the basic feature extraction layer adopting resnet101 as an example.
[0130] As shown in Table 1, it is the structural parameters of the basic feature extraction layer with resnet101 as an example.
[0131] Among them, the size of the feature map output by convolutional layer 1 is 300x500, the size of the convolutional kernel of convolutional layer 1 is 7x7, the number of channels is 64, and the convolutional stride (stride) is 2; the size of the feature map finally output by convolutional layer 2 is 150x250. Convolutional layer 2 includes a max - pooling layer with a size of 3x3 and a stride of 2, and also includes 3 types of convolutional kernels, namely a convolutional kernel with a size of 1x1 and a channel number of 64, a convolutional kernel with a size of 3x3 and a channel number of 64, and a convolutional kernel with a size of 1x1 and a channel number of 256. The subsequent convolutional layers 3 / 4 and 5 are similar by analogy.
[0132]
[0133] Table 1
[0134] In the embodiment of the present application, the parameter initialization of the above - mentioned basic feature extraction layer, such as Conv1 - Conv5 in Table 1 above, can adopt the model parameters of the Imagenet pre - trained model, that is, the parameters of ResNet101 pre - trained on the ImageNet dataset. Of course, the model parameters of the openimage pre - trained model can also be adopted, that is, the basic feature extraction model pre - trained on the openimage dataset.
[0135] In the embodiments of the present application, the feature compression layer is used to compress the sparse high-risk vectors extracted by the basic feature extraction layer into dense low-dimensional vectors. Since the computer memory space is limited, for large-scale retrieval, the more compact the features, the smaller the storage space. Therefore, at the same time (or under the same memory), a larger retrieval library can be accommodated for similarity comparison. Therefore, the purpose of feature compression can not only make the features denser, but also reduce the feature storage space and improve the subsequent vector retrieval efficiency. For example, when compressing a 1*2048-dimensional feature vector output by the basic feature extraction layer into a dense 1*128-dimensional feature vector, if storing a 1*2048-dimensional vector requires 32 bytes (B), while a 1*128-dimensional vector only requires 2 bytes for storage. For 100 million feature vectors, if not compressed, it requires 1525.9 MB, that is, 1.5 GB, while after compression, it only requires 2 million B, that is, 95.4 MB. Obviously, the storage space utilization rate is higher.
[0136] See Figure 5 As shown, the feature compression layer may include a pooling layer and two fully connected (fc) layers. As shown in Table 2, it is a schematic diagram of the structural parameters of the feature compression layer.
[0137] Layer Name Output Feature Size Layer Operation Pooling Layer 1x2048 Max Pooling Fully Connected Layer 1 1x512 Fully Connected Fully Connected Layer 2 1x128 Fully Connected
[0138] Table 2
[0139] Among them, the pooling layer is used to perform pooling processing on the basic feature vectors output by the basic feature extraction layer. As shown in Table 2, a feature vector of size 1x2048 can be obtained. The pooling layer can use max pooling as shown in Table 2. Of course, other possible pooling methods can also be used, such as average pooling, etc. The first fully connected layer and the second fully connected layer respectively perform two compressions on the feature vectors after pooling to obtain feature representation vectors with lower dimensions. As shown in Table 2, the feature vector obtained by the first fully connected layer has a size of 1x512, and the feature vector obtained by the second fully connected layer has a size of 1x128.
[0140] For the parameter initialization of the above feature compression layer, the fully connected layer can be initialized, for example, using a Gaussian distribution with a variance of 0.01 and a mean of 0.
[0141] It should be noted that the above model structure and parameters are only examples of one possibility. In actual applications, other possible structures and parameters can also be used, and the embodiments of the present application do not limit this.
[0142] Step 403: Determine the feature representation loss value.
[0143] In the embodiments of the present application, see Figure 3 As shown, after obtaining the feature representation vectors of each image sample, the feature representation loss value of the feature encoding sub-model can be determined based on the feature representation vectors.
[0144] In a possible implementation, since each image sample group of the structure is labeled, the feature representation loss value can be determined based on the labels of each image sample group.
[0145] Specifically, for the feature representation vectors of at least two image samples included in each of the multiple image sample groups, the predicted similarity between at least two image samples can be obtained, and the obtained predicted similarity can be compared with the corresponding labeled similarity to obtain the comparison result corresponding to each image sample group. Furthermore, the feature representation loss value can be determined according to the respective comparison results obtained for each image sample group.
[0146] Specifically, taking an image sample group including two image samples as an example, the feature representation loss value can be expressed in the following manner.
[0147]
[0148] Among them, L1 represents the feature representation loss value of the feature encoding sub-model, S i1 represents the predicted similarity between two image samples in the i-th image sample group, and the predicted similarity can be calculated based on the feature representation vectors of the two image samples in the i-th image sample group. S i2 represents the labeled similarity between two image samples in the i-th image sample group, and N represents the number of image sample groups.
[0149] In another possible implementation, each image sample group can be a triple. Then, an image sample group can include a first image sample (or called anchor image), a second image sample (or called positive image), and a third image sample (or called negative image). And the similarity between the first image sample and the second image sample in each image sample group is not less than a preset similarity threshold, and the similarity between the first image sample and the third image sample is less than the similarity threshold. That is, among the three image samples, the first image sample and the second image sample are similar samples, while the first image sample and the third image sample are dissimilar samples.
[0150] Specifically, when each image sample group can be a triple, when determining the feature representation loss value, for each image sample group, the first similarity between the first image sample and the second image sample in each image sample group can be obtained respectively, and the second similarity between the first image sample and the third image sample can be obtained. And based on the first similarity and the second similarity, the triple loss value corresponding to an image sample group can be obtained. Furthermore, based on the respective triple loss values obtained for each image sample group, the feature representation loss value can be obtained.
[0151] The feature representation loss value can be expressed in the following way:
[0152]
[0153] Among them, L1 represents the feature representation loss value of the feature encoding sub-model, and d i (a, p) represents the distance between the first image sample and the second image sample in the i-th group of image samples, and d i (a, n) represents the distance between the first image sample and the third image sample in the i-th group of image samples. margin is a constant greater than 0, and N represents the number of groups of image samples.
[0154] Of course, according to different sample constructions, an appropriate method can be adopted to obtain the feature representation loss value.
[0155] Step 404: Cluster multiple feature representation vectors and determine the clustering loss value based on the clustering result.
[0156] See Figure 3 As shown, after obtaining the feature representation vectors corresponding to multiple image samples, in addition to calculating the corresponding feature representation loss value, clustering is also performed based on the feature representation vectors corresponding to multiple image samples to obtain a clustering result.
[0157] Specifically, the clustering sub-model is implemented through a clustering layer. See Figure 3 As shown, it is a schematic diagram of the structural parameters of the clustering layer. The role of the clustering layer is to perform feature projection on the feature representation vectors of each image sample. The clustering layer stores the clustering center vectors of each category, that is, maps the 1*128 feature vector (1*128 is an example of the feature representation vector corresponding to Table 2 above) to N category centers, where N is the number of categories. For example, for a clustering task of 100,000 categories, N = 100,000, then the parameters of the clustering layer can be 128*100,000, which consists of 100,000 category center vectors of 128*1. The clustering layer maps the embedding of the image 1*128 to a certain category in the 100,000 clusters through vector similarity calculation.
[0158] Layer Name Output Size Layer Operation Clustering Layer 1xN Fully Connected
[0159] Table 3
[0160] It should be noted that the clustering layer shown in Table 3 only contains a structure of a single fully connected layer, which is only one possible example. In practical applications, other clustering layer structures can also be adopted. For example, one possible way is to deepen the clustering layer structure in Table 3, such as adding multiple fully connected layers, rectified linear unit (ReLU) activation layers, or convolutional layers, etc. Another possible way is to adopt other deep neural network structures.
[0161] Therefore, the essence of the clustering task is to learn the clustering center vector. In the embodiments of the present application, it is to learn by converting the clustering center vector into network weight parameters of deep learning. Through multiple rounds of clustering processes, the clustering center vector is continuously updated. When the image clustering model converges, the final clustering center vector is obtained.
[0162] In a possible implementation manner, when using, for example, triplet loss for feature learning, in each round of clustering, only image samples of different categories in each triplet can be selected to participate in the clustering. For example, the first image sample and the third image sample can be selected, or the second image sample and the third image sample can be selected.
[0163] Taking the selection of the first image sample and the third image sample as an example, when clustering, at least one feature representation vector corresponding to the first image sample and the third image sample in each group of image samples is selected from the obtained multiple feature representation vectors, and then at least one feature representation vector is clustered.
[0164] As for the category of the second image sample, it can be assigned to the category to which the first image sample belongs, or after clustering, the similarity between the second image sample and each clustering center vector can be calculated, and then the category with the largest similarity is determined as the category to which the second image sample belongs.
[0165] In a possible implementation manner, when using, for example, triplet loss for feature learning, in each round of clustering, after clustering, image samples of different categories in each triplet can be selected to participate in the calculation of the clustering loss.
[0166] Specifically, for each feature representation vector among the multiple feature representation vectors, it is respectively mapped to each image category to obtain the corresponding image category vector. See Figure 6 As shown, it is a schematic diagram of an image category vector, where the image category vector contains multiple elements ( Figure 6 each box in represents an element position), and each element corresponds to an image category. The value of each element represents whether the image sample corresponding to a feature representation vector belongs to the corresponding image category. For example, Figure 6As shown, if the value of the first element position is 1, it indicates that the image sample belongs to the image category corresponding to the first element position. If the value of the third element position is 0, it indicates that the image sample does not belong to the image category corresponding to the third element position. Of course, other values can also be used to represent whether it belongs to the corresponding image category. For example, 0 can be used to represent belonging to the corresponding image category, and 1 can be used to represent not belonging to the corresponding image category. The embodiments of the present application do not limit this.
[0167] Taking a feature representation vector A as an example, when obtaining its image category vector, the similarities between the feature representation vector A and each clustering center vector determined in this round of clustering are respectively determined. Based on the obtained similarities, it is respectively determined whether the image sample corresponding to the feature representation vector A belongs to the image category represented by the corresponding clustering center vector, and a determination result is obtained. Then, based on the determination result, the image category vector is obtained.
[0168] For example, when calculating the similarity between the feature representation vector A and the clustering center vector of category B, if the similarity between the feature representation vector A and the clustering center vector of category B is greater than the set threshold, it is determined that the image sample corresponding to the feature representation vector A belongs to category B. Then, in the image category vector, the value of the corresponding element position of category B is the value representing belonging to this category. By calculating and judging the similarities of each category, the image category vector of the feature representation vector A can be obtained.
[0169] Alternatively, after calculating the similarities between the feature representation vector A and the clustering center vectors of each category respectively, the one with the largest similarity is selected as the category to which the image sample corresponding to the feature representation vector A belongs. Then, the value of the corresponding element position in the image category vector is set to the value representing belonging to this category, thereby obtaining the image category vector of the feature representation vector A.
[0170] According to the obtained image category vector, the image categories to which each image sample belongs after this round of clustering can be determined. Thus, the image category vector is equivalent to the label corresponding to each image sample. Since this label is generated by clustering and is not the true label of the image sample, it is called the clustering pseudo-label here. When the clustering sub-model tends to converge, the clustering pseudo-labels of each image sample no longer change. Therefore, the change of the clustering pseudo-label can be used to measure whether the clustering sub-model converges.
[0171] Furthermore, compare the image class vectors obtained in this round with the image class vectors determined for the image samples corresponding to a feature representation vector in the previous round of clustering to obtain a comparison result, that is, compare the image class vectors obtained in this round with the image class vectors of the previous round for the same image sample to obtain the comparison results corresponding to the feature representation vectors of each image sample, and determine the cluster loss value based on the multiple comparison results obtained for multiple feature representation vectors.
[0172] For example, the cluster loss value can be expressed as follows:
[0173]
[0174] where L2 represents the feature representation loss value of the clustering sub-model, R j1 represents the image class vector obtained for the j-th image sample in this round of clustering, and R j2 represents the image class vector obtained for the j-th image sample in the previous round.
[0175] In the embodiments of the present application, when using, for example, triplet loss for feature learning, during each round of clustering, only the image samples of different classes in each triplet can be selected to participate in the calculation of the cluster loss. After clustering all the image samples, select the image samples of different classes in each triplet to participate in the calculation of the cluster loss. For example, the first image sample and the third image sample can be selected, or the second image sample and the third image sample can be selected.
[0176] Taking the selection of the first image sample and the third image sample as an example, after clustering, select the comparison results corresponding to at least one feature representation vector of the first image sample and the third image sample in each group of image samples to calculate the cluster loss value.
[0177] In the embodiments of the present application, in order to ensure that the learned by the clustering layer is the clustering center, it is necessary to perform L2 normalization on the embedding to ensure that the length of the embedding vector is 1. At the same time, after each update of the parameters of the clustering layer, the length of each clustering center vector of the clustering layer should also be 1.
[0178] Step 405: Determine the total model loss value of the image clustering model based on the feature representation loss value and the cluster loss value.
[0179] In the embodiments of the present application, all parameters of the model are set to the state to be learned. During training, the neural network performs forward calculations on each input image sample to obtain prediction results, including the prediction results in the feature learning stage and the prediction results in the clustering stage. And the corresponding loss values, namely the feature representation loss value and the clustering loss value, are obtained based on the prediction results of each stage respectively. Furthermore, the total model loss value of the image clustering model for this round of clustering is obtained based on the feature representation loss value and the clustering loss value.
[0180] In the embodiments of the present application, the clustering task updates the image class vector of each image sample. In the next round of learning, the learning target of the clustering layer is the updated image class vector (i.e., the clustering target shown in Figure 3 . However, since the clustering loss value will increase after each clustering task updates the image class vector, which affects the feature encoding process of the model through gradient backpropagation. Therefore, during the model training process, different weights are designed for the influence degree of clustering in different stages. And the clustering weight is adjusted according to the situation of feature learning. That is, when there are large changes or oscillations in the feature representation loss value, it indicates that the features are not converged or a certain clustering brings changes to the embedding. At this time, it is necessary to adjust the embedding well first and then perform clustering. Therefore, it is necessary to dynamically adjust the clustering loss weight. Thus, the weight value of the clustering loss value can be determined based on the difference between the feature representation loss value and the feature representation loss value of the previous round of clustering. Furthermore, based on the feature representation loss value, the clustering loss value, and the weight value, the total model loss of the image clustering model is determined, where the weight value is negatively correlated with the difference.
[0181] Since the clustering loss and the feature representation loss are usually not in the same order of magnitude, before determining the total model loss value of the image clustering model based on the feature representation loss value and the clustering loss value, the feature representation loss value and the clustering loss value can also be subjected to scale transformation to make the measurement scales of the feature representation loss value and the clustering loss value the same, achieving magnitude balance.
[0182] The total model loss value can be obtained through the following formula:
[0183] L = L1 + b * L2
[0184]
[0185] Wherein, L1 represents the feature representation loss value obtained in this round, L1' represents the feature representation loss value of the previous round, abs(x) represents taking the absolute value, and scale represents the scale adjustment degree of the clustering loss. Since the loss of clustering is often 10 times the order of magnitude of feature learning, scale can be taken as 0.1 to achieve magnitude balance through scale.
[0186] In the early stage of model learning, the embedding has not converged. After the clustering task is updated, the change in the clustering loss is obvious. At this time, due to the non-convergence of the embedding, the change in the feature representation loss value is very large. In the total loss value L of the model, the adjustment weight b is relatively small. For example, if the feature representation loss value changes from 7 to 2, then b is 1 / ((7 - 2) / 2)*0.1 = 0.04, and b is relatively small. Therefore, even if the clustering task changes drastically at this time, the impact on the learning of the embedding is not very large. In the later stage of model learning, when the clustering task is updated and the change in the clustering loss is obvious, but at this time the embedding has converged and the change in the feature representation loss value is not large. Therefore, the loss adjustment weight becomes larger. For example, if the feature representation loss value changes from 0.6 to 0.56, then b is min(1, 1 / ((0.6 - 0.56) / 0.6)*0.1) = 1, and the model fully learns the clustering, and the maximum loss weight b is 1.
[0187] Step 406: Determine whether the image clustering model has converged based on the total loss value of the model.
[0188] Step 407: If the determination result in Step 406 is yes, stop the clustering iteration.
[0189] Step 408: If the determination result in Step 406 is no, update the parameters of the image clustering model.
[0190] In the embodiments of the present application, it is possible to determine whether the image clustering model has converged according to whether the total loss value of the model is less than the set loss value. When the total loss value of the model is less than the set loss value, the image clustering model has converged, and then the clustering iteration process can be stopped. If the total loss value of the model is not less than the set loss value, the image clustering model has not converged, and then the adjustment gradient of the model parameters can be calculated according to the total loss value of the model, and further the parameters of the image clustering model can be adjusted, and the adjusted image clustering model is used to enter the next round of clustering.
[0191] Next, a specific example is used to introduce the process of determining the clustering center.
[0192] In one implementation, the set of image samples to be clustered includes multiple image samples. The image samples can be images containing animals such as cats, dogs, chickens, or ducks, or images containing plants such as cherry trees, pear trees, rice, or wheat. Of course, they can also be images containing other content. After clustering by the clustering center determination method provided by the embodiments of the present application, multiple image samples can be divided into multiple clusters, and each cluster is an image of the same category. For example, the image samples with chickens in the images are divided into the same cluster, and the image samples with dogs in the images are divided into the same cluster. Then, the feature representation vector of the image sample that can best represent the cluster is selected from each cluster as the clustering center vector of the cluster.
[0193] In the embodiments of the present application, based on the above description, each time the image category vector is updated in the clustering task, the clustering loss value will increase, thereby affecting the feature encoding process of the model through gradient backpropagation. Therefore, considering the efficiency of model training, the clustering task update cannot be too frequent. Therefore, after multiple rounds of update iterations of the feature encoding sub-model, the clustering task can be updated once. For example, the clustering task can be updated once every 5 rounds of model iterations. Therefore, referring to Figure 7 FIG. 3 is another schematic diagram of the clustering process provided by the embodiments of the present application.
[0194] Step 701: Construct multiple groups of image samples based on the image sample set.
[0195] Step 702: Perform feature encoding on each image sample respectively, and obtain corresponding multiple feature representation vectors based on the encoding results.
[0196] Step 703: Determine the feature representation loss value based on the obtained multiple feature representation vectors.
[0197] Step 704: Determine whether the current round of clustering meets the set conditions.
[0198] Among them, the set conditions may include that the number of times of feature encoding in the current round of clustering reaches the set number threshold. For example, when the clustering task is updated once every 5 rounds of model iterations as described above, when the number of times of feature encoding reaches 5 times, the set conditions are met. Or, the feature encoding sub-model meets the convergence conditions.
[0199] Step 705: If the determination result of step 704 is no, adjust the parameters of the feature encoding sub-model, and return to step 702.
[0200] If the determination result of step 704 is yes, jump to step 706.
[0201] Step 706: Perform clustering based on the multiple feature representation vectors obtained in the last feature encoding in the current round of clustering, and determine the clustering loss value based on the clustering result.
[0202] Step 707: Determine the total loss value of the image clustering model based on the feature representation loss value and the clustering loss value.
[0203] Step 708: Determine whether the image clustering model converges based on the total loss value of the model.
[0204] Step 709: If the determination result of step 708 is yes, stop the clustering iteration.
[0205] Step 710: If the determination result of step 708 is no, update the parameters of the image clustering model.
[0206] Figure 7 In Figure 4 the similar steps in Figure 4 the relevant part can be referred to, and will not be elaborated here.
[0207] In the embodiment of the present application, after the image clustering model converges, the clustering center vectors can be obtained, that is, the multiple clustering center vectors included in the clustering result obtained in the last round of clustering.
[0208] Specifically, the feature representation vectors that can represent each category in the clustering result can be used as the clustering center vectors for each category, whereby multiple clustering center vectors can be obtained. At the same time, the feature representation vectors of each image sample can also be obtained.
[0209] The embedding online (i.e., learning while clustering) clustering method provided in the embodiment of the present application uses the features for clustering obtained from the model adjusted based on the feedback of the clustering process. Therefore, the embeddings of each image sample have, to a certain extent, the distribution characteristics of the previous kmeans clustering. Continuing to cluster on this basis can maintain the similarity between the new clustering and the previous clustering to a greater extent, that is, if two image samples are clustered into the same label this time, they are very likely to be in the same label after the next clustering rather than different labels. The number of samples with this category retention property increases as the embedding becomes more stable.
[0210] In the embodiment of the present application, the obtained clustering center vectors and feature representation vectors can be used in subsequent image retrieval or video recommendation applications. Refer to Figure 8 and Figure 9 as shown, Figure 8 which is a schematic diagram of the application process taking the video recommendation scenario as an example, Figure 9 and is a schematic diagram of video retrieval.
[0211] Step 801: For each video in the video database, based on multiple clustering center vectors, respectively determine the image categories to which each image frame included in each video belongs, and associate each video with each image category according to the determined image categories.
[0212] In the embodiment of the present application, when training the image clustering model, the image frames of each video can be used as image samples for training. Furthermore, after the training is completed, the clustering center vectors and the feature representation vectors of the image frames of each video can be obtained simultaneously. Refer to Figure 9 as shown, the video library includes N videos, namely Video A to Video N, and each video corresponds to at least one image frame, such as Figure 9The video A shown contains images A1 to An, thus forming an image library. After the image clustering model converges, the image representation vector of each image frame can be obtained, and then the similarity between each image frame and each cluster center vector is calculated, and the image category corresponding to the most similar cluster center vector is determined as the image category of each image frame, and then the video in which it is located is associated with the corresponding image category. For example, when the most similar cluster center vector of image A1 is the center vector of category 1, the image category of image A1 is category 1, and video A is associated with category 1. In this way, each category 1 can be associated with at least one video to complete the bucketing of the video.
[0213] Step 802: Receive a video retrieval request sent by a target user, where the video retrieval request carries a target video.
[0214] When the user triggers video retrieval on the video retrieval interface, the terminal device initiates a video retrieval request, and the video retrieval request carries the target video to be retrieved.
[0215] Step 803: Based on multiple cluster center vectors, determine the image categories to which each image frame included in the target video belongs, and determine the videos associated with each determined image category as candidate matching videos.
[0216] In the embodiment of the present application, after acquiring the target video, each image frame included in the target video can be Figure 9 The target images 1 to N shown in the figure respectively obtain their corresponding image category vectors through the image clustering model provided in the embodiment of the present application, that is, feature encoding is performed on them to obtain their feature representation vectors, and the image category to which they belong is determined based on the feature representation vectors and each cluster center vector, and the videos associated with the corresponding image category are recalled as candidate matching videos for the target video.
[0217] Taking target image 1 as an example, feature encoding is performed on target image 1 to obtain its feature representation vector 1, and then the feature representation vector 1 is respectively calculated with each cluster center vector for similarity, and the image category corresponding to the cluster center vector with the greatest similarity is determined as the image category to which target image 1 belongs. For example, the image category to which target image 1 belongs is category 1, and then all videos associated with category 1 are recalled as candidate matching videos for the target video. Of course, for each image frame, the K categories with the highest similarity can also be selected, and all videos associated with these K categories can be recalled.
[0218] Since, in actual applications, each video includes a large number of image frames, some image frames can be selected therefrom to perform the process of video bucket recall. In one possible implementation, key frames can be selected to perform the above process for video bucket recall. In one possible implementation, the similarities between all the image frames of the target video and each cluster center vector can be obtained, the similarities of each image frame to its category can be obtained, and then all the videos associated with the categories corresponding to the top K similarities with the highest similarity are recalled as the candidate matching videos of the target video.
[0219] Step 804: Select a target matching video to be recommended to the target user from the candidate matching videos.
[0220] For the recalled videos above, refined matching calculation for retrieval and final similarity need to be performed to determine the target matching video to be recommended to the target user from the candidate matching videos. Among them, for retrieval refined matching, generally Scale-invariant feature transform (SIFT) image features can be used, that is, first calculate the SIFT features of two samples respectively, and then perform matching according to the SIFT features. If the similarity is greater than the threshold, it means the two samples are similar.
[0221] In summary, in the embodiments of the present application, end-to-end feature learning and clustering are realized by means of the multi-task combination of deep learning image embedding and clustering, and dynamic loss adjustment is designed to ensure the stability of feature embedding. Finally, the learned embedding has a better effect in similarity clustering. Specifically, the clustering similarity is improved by jointly learning a deep neural network of embedding and unsupervised clustering, the problem of insufficient global representation of metric learning is reduced by global similarity clustering, and the stability of embedding convergence under multiple tasks is realized through dynamic loss feedback and adjustment, and end-to-end feature learning and clustering are realized under the original labeled data. For example, triplet labeled triples can be used to train the embedding, and the clustering uses an unsupervised method to generate learning objectives, optimizing the clustering semantic similarity problem at the feature level without introducing additional labeled data.
[0222] Moreover, since, through the clustering center determination method of the embodiments of the present application, the image samples clustered into the same category are closer to each other, the similarity of the effect of bucket recall is better than that of the two-stage feature and clustering split learning.
[0223] Please refer to Figure 10 , based on the same inventive concept, the embodiments of the present application also provide a clustering center determination device 100, and the device includes:
[0224] A sample acquisition unit 1001, configured to acquire a set of image samples;
[0225] A clustering unit 1002, configured to perform multiple rounds of clustering on the set of image samples by using an image clustering model until a preset convergence condition is satisfied, where the clustering unit includes a feature encoding subunit 10021, a clustering subunit 10022, and a model adjustment subunit 10023:
[0226] The feature encoding subunit 10021 is configured to respectively perform feature encoding on each image sample in the set of image samples, obtain corresponding multiple feature representation vectors based on the encoding results, and determine a feature representation loss value;
[0227] The clustering subunit 10022 is configured to cluster the multiple feature representation vectors and determine a clustering loss value based on the clustering results;
[0228] When it is determined that the image clustering model has not converged based on the feature representation loss value and the clustering loss value, the model adjustment subunit 10023 is configured to adjust the parameters of the image clustering model;
[0229] An output unit 1003, configured to output multiple cluster center vectors included in the clustering results obtained in the last round of clustering, where one cluster center vector corresponds to one image category.
[0230] Optionally, the feature encoding subunit 10021 is specifically configured to:
[0231] For each image sample, perform at least one time of basic feature extraction respectively to obtain corresponding basic feature vectors;
[0232] Respectively perform feature compression on each obtained basic feature vector to obtain a feature representation vector corresponding to each image sample.
[0233] Optionally, the sample acquisition unit 1001 is further configured to:
[0234] Construct multiple groups of image samples based on the set of image samples, where each group of image samples includes at least two image samples, and includes the labeled similarity between at least two image samples;
[0235] Then the feature encoding subunit 10021 is specifically configured to:
[0236] For multiple groups of image samples, respectively perform the following operations: based on the feature representation vectors of at least two image samples included in one group of image samples in the multiple groups of image samples, obtain the predicted similarity between at least two image samples, and compare the obtained predicted similarity with the corresponding labeled similarity to obtain a corresponding comparison result;
[0237] Determine the feature representation loss value according to each comparison result obtained for each group of image samples.
[0238] Optionally, each group of image samples includes a first image sample, a second image sample, and a third image sample, and the similarity between the first image sample and the second image sample in each group of image samples is not less than a preset similarity threshold, and the similarity between the first image sample and the third image sample is less than the similarity threshold;
[0239] Then, the feature encoding subunit 10021 is specifically configured to:
[0240] For multiple groups of image samples, perform the following operations respectively: In one group of image samples within at least one group of image samples, obtain the first similarity between the first image sample and the second image sample, and obtain the second similarity between the first image sample and the third image sample, and obtain a triplet loss value corresponding to one group of image samples based on the first similarity and the second similarity;
[0241] Obtain the feature representation loss value based on the triplet loss values corresponding to each group of image samples obtained.
[0242] Optionally, the clustering subunit 10022 is specifically configured to:
[0243] Select at least one feature representation vector corresponding to the first image sample and the third image sample in multiple groups of image samples from multiple feature representation vectors;
[0244] Perform clustering on at least one feature representation vector.
[0245] Optionally, the clustering subunit 10022 is specifically configured to:
[0246] For multiple feature representation vectors, perform the following operations respectively:
[0247] Map one feature representation vector among multiple feature representation vectors to each image category respectively to obtain corresponding image category vectors; wherein, each element in the image category vector corresponds to an image category, and the value of each element represents whether the image sample corresponding to a feature representation vector belongs to the corresponding image category;
[0248] Compare the image category vector with the image category vector determined for the image sample corresponding to one feature representation vector in the previous round of clustering to obtain a comparison result;
[0249] Determine the clustering loss value based on the multiple comparison results obtained for multiple feature representation vectors.
[0250] Optionally, the clustering subunit 10022 is specifically configured to:
[0251] Select at least one feature representation vector corresponding to the first image sample and the third image sample from multiple feature representation vectors in multiple groups of image samples;
[0252] Determine a clustering loss value based on multiple comparison results obtained for multiple feature representation vectors, including:
[0253] Determine the clustering loss value based on the comparison result corresponding to at least one feature representation vector.
[0254] Optionally, the clustering subunit 10022 is specifically configured to:
[0255] Respectively determine the similarity between a feature representation vector and each clustering center vector determined in this round of clustering;
[0256] Based on the obtained similarities, respectively determine whether the image sample corresponding to a feature representation vector belongs to the image category represented by the corresponding clustering center vector, and obtain a determination result;
[0257] Based on the determination result, obtain an image category vector.
[0258] Optionally, the model adjustment subunit 10023 is further configured to:
[0259] Determine a weight value of the clustering loss value based on the difference between the feature representation loss value and the feature representation loss value of the previous round of clustering; wherein, the weight value is negatively correlated with the difference;
[0260] Determine the total model loss value of the image clustering model based on the feature representation loss value, the clustering loss value, and the weight value;
[0261] Determine whether the image clustering model converges based on the total model loss value.
[0262] Optionally, the model adjustment subunit 10023 is further configured to:
[0263] Perform scale transformation on the feature representation loss value and the clustering loss value so that the measurement scales of the feature representation loss value and the clustering loss value are the same.
[0264] Optionally, the feature encoding subunit 10021 is further configured to:
[0265] Perform at least one iterative update on the adjusted feature encoding sub-model until the current round of clustering meets the set conditions; wherein, in each iterative update, perform feature encoding on each image sample based on the feature encoding sub-model, determine a feature representation loss value based on the multiple feature representation vectors obtained by encoding, and adjust the parameters of the feature encoding sub-model based on the feature representation loss value, and the set conditions include that the number of times of feature encoding in the current round of clustering reaches a set number threshold, or the feature encoding sub-model meets the convergence condition;
[0266] Then the clustering sub-unit 10022 is specifically used for:
[0267] Perform clustering based on the multiple feature representation vectors obtained by encoding using the feature encoding sub-model after the last iterative update in the current round of clustering.
[0268] Optionally, the device further includes a video bucketing unit 1004 and a video retrieval unit 1005,
[0269] The video bucketing unit 1004 is used to perform the following operations respectively for each video in the video database: based on multiple clustering center vectors, respectively determine the image categories to which each image frame included in a video belongs, and associate a video to each image category according to the determined image categories;
[0270] The video retrieval unit 1005 is used to receive a video retrieval request sent by a target user, and the video retrieval request carries a target video; based on multiple clustering center vectors, determine the image categories to which each image frame included in the target video belongs, and determine the videos associated with the determined image categories as candidate matching videos; select a target matching video to recommend to the target user from the candidate matching videos.
[0271] The device can be used to execute Figures 2 - 9 the method shown in the embodiments shown, therefore, for the functions that can be realized by each functional module of the device, etc., reference can be made to the description of the embodiments shown in Figures 2 - 9 shown, and details will not be repeated. Among them, the video bucketing unit 1004 and the video retrieval unit 1005 are not essential functional modules, so they are shown in dashed lines in Figure 10 the figure.
[0272] The device can adopt the constructed end-to-end image clustering model, which combines the feature learning and the clustering process, that is, unsupervised clustering is carried out while feature learning is in progress, so that the feature learning and the clustering promote and optimize each other. In other words, during the feature learning process, to a certain extent, the distribution characteristics of the previous clustering can be learned, and then the learned feature representation vectors will also tend to certain distribution characteristics, improving the accuracy of feature learning. In addition, based on this, continuing to cluster can maintain the similarity between the new clustering and the previous clustering to a greater extent, making the features clustered into the same category closer to each other. Correspondingly, the finally obtained clustering centers will be more accurate, and thus the clustering effect will be better. For example, when the device is used for video bucketing retrieval, due to the better clustering effect, it lays a foundation for the accuracy of subsequent video bucketing retrieval.
[0273] Please refer to Figure 11 , based on the same technical concept, the embodiment of the present application also provides a computer device 110, which may include a memory 1101 and a processor 1102. The computer device 110 may be a terminal or a server, for example.
[0274] The memory 1101 is used to store the computer program executed by the processor 1102. The memory 1101 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the computer device, etc. The processor 1102 may be a central processing unit (CPU) or a digital processing unit, etc. In the embodiment of the present application, the specific connection medium between the above-mentioned memory 1101 and the processor 1102 is not limited. In the embodiment of the present application Figure 11 it is connected by a bus 1103 between the memory 1101 and the processor 1102. The bus 1103 is represented by a thick line in Figure 11 For the connection manners of other components, only a schematic illustration is given and is not limited thereto. The bus 1103 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only a thick line is used to represent it in
[0275] The memory 1101 may be a volatile memory, such as a random-access memory (RAM); the memory 1101 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the memory 1101 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1101 may be a combination of the above memories.
[0276] The processor 1102 is configured to execute the method performed by the device in the embodiment as Figures 2 - 9 shown when calling the computer program stored in the memory 1101.
[0277] In some possible implementation manners, each aspect of the method provided in this application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the method according to various exemplary implementation manners of this application described above in this specification. For example, the computer device may execute the method performed by the device in the embodiment as Figures 2 - 9 shown.
[0278] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0279] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of this application.
[0280] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. A method for determining a clustering center, characterized in that The method includes: Obtaining a set of image samples, and constructing a plurality of groups of image samples based on the set of image samples, each group of image samples including at least two image samples; Performing multiple rounds of clustering using an image clustering model until a preset convergence condition is met, where one round of clustering includes the following operations: Performing feature encoding on each of the image samples included in the plurality of groups of image samples, and obtaining a corresponding plurality of feature representation vectors based on the encoding results; Determining a feature representation loss value according to the similarity between at least two image samples included in each of the plurality of groups of image samples; Clustering the plurality of feature representation vectors, and determining a clustering loss value based on the clustering results; When it is determined that the image clustering model has not converged based on the feature representation loss value and the clustering loss value, adjusting the parameters of the image clustering model; Outputting a plurality of cluster center vectors included in the clustering results obtained in the last round of clustering, with one cluster center vector corresponding to one image category.
2. The method according to claim 1, wherein Performing feature encoding on each of the image samples in the set of image samples, and obtaining a corresponding plurality of feature representation vectors based on the encoding results, includes: Performing at least one basic feature extraction on each of the image samples to obtain corresponding basic feature vectors; Performing feature compression on each of the obtained basic feature vectors to obtain the feature representation vectors corresponding to each of the image samples.
3. The method according to claim 1, characterized in that Each group of image samples includes the labeled similarity between the at least two image samples; Then, determining the feature representation loss value according to the similarity between at least two image samples included in each of the plurality of groups of image samples, includes: For each of the plurality of groups of image samples, respectively perform the following operations: Based on the feature representation vectors of at least two image samples included in one group of image samples within the plurality of groups of image samples, obtain the predicted similarity between the at least two image samples, and compare the obtained predicted similarity with the corresponding labeled similarity to obtain the corresponding comparison result; Determine the feature representation loss value according to the respective comparison results obtained for each group of image samples.
4. The method according to claim 1, wherein The at least two image samples include a first image sample, a second image sample, and a third image sample, and the similarity between the first image sample and the second image sample in each group of image samples is not less than a preset similarity threshold, and the similarity between the first image sample and the third image sample is less than the similarity threshold; Then, determining the feature representation loss value according to the similarity between at least two image samples included in each of the plurality of groups of image samples, includes: For each of the plurality of groups of image samples, respectively perform the following operations: In one group of image samples within the plurality of groups of image samples, obtain the first similarity between the first image sample and the second image sample, and obtain the second similarity between the first image sample and the third image sample, and based on the first similarity and the second similarity, obtain the triplet loss value corresponding to the one group of image samples; Obtain the feature representation loss value based on the respective triplet loss values corresponding to each group of image samples obtained.
5. The method according to claim 4, wherein Clustering the multiple feature representation vectors includes: Selecting at least one feature representation vector corresponding to the first image sample and the third image sample in the multiple image sample groups from the multiple feature representation vectors; Clustering the at least one feature representation vector.
6. The method according to claim 4, wherein Clustering the multiple feature representation vectors and determining a clustering loss value according to the clustering result, including: Performing the following operations respectively on the multiple feature representation vectors: Mapping one feature representation vector in the multiple feature representation vectors to each image category respectively to obtain corresponding image category vectors; wherein, each element in the image category vector corresponds to an image category, and the value of each element represents whether the image sample corresponding to the one feature representation vector belongs to the corresponding image category; Comparing the image category vector with the image category vector determined for the image sample corresponding to the one feature representation vector in the previous round of clustering to obtain a comparison result; Determining the clustering loss value based on the multiple comparison results obtained for the multiple feature representation vectors.
7. The method according to claim 6, wherein Before determining the clustering loss value based on the multiple comparison results obtained for the multiple feature representation vectors, the method further includes: Selecting at least one feature representation vector corresponding to the first image sample and the third image sample in the multiple image sample groups from the multiple feature representation vectors; Determining the clustering loss value based on the multiple comparison results obtained for the multiple feature representation vectors, including: Determining the clustering loss value based on the comparison results corresponding to the at least one feature representation vector.
8. The method according to claim 6, wherein Mapping one feature representation vector in the multiple feature representation vectors to each image category respectively to obtain corresponding image category vectors, including: Determining the similarity between the one feature representation vector and each clustering center vector determined in this round of clustering respectively; Based on the obtained similarities, determining whether the image sample corresponding to the one feature representation vector belongs to the image category represented by the corresponding clustering center vector respectively to obtain a determination result; Obtaining the image category vector based on the determination result.
9. The method according to any one of claims 1-8, characterized in that, Before adjusting the parameters of the image clustering model when it is determined that the image clustering model has not converged based on the feature representation loss value and the clustering loss value, the method further includes: Determining a weight value of the clustering loss value based on the difference between the feature representation loss value and the feature representation loss value of the previous round of clustering; wherein, the weight value is negatively correlated with the difference; Determining a total model loss value of the image clustering model based on the feature representation loss value, the clustering loss value, and the weight value; Determining whether the image clustering model converges based on the total model loss value.
10. The method according to claim 9, wherein Before determining the total model loss value of the image clustering model based on the feature representation loss value, the clustering loss value, and the weight value, the method further includes: Performing a scale transformation on the feature representation loss value and the clustering loss value so that the measurement scales of the feature representation loss value and the clustering loss value are the same.
11. The method according to any one of claims 1-8, characterized in that, The image clustering model includes a feature encoding sub-model. Before clustering the multiple feature representation vectors and determining the clustering loss value based on the clustering result, the method further includes: When it is determined that the feature encoding sub-model has not converged based on the feature representation loss value, adjusting the parameters of the feature encoding sub-model; Performing at least one iterative update on the adjusted feature encoding sub-model until the current round of clustering meets the set conditions; wherein, in each iterative update, respectively performing feature encoding on each image sample based on the feature encoding sub-model, determining the feature representation loss value based on the multiple feature representation vectors obtained by encoding, and adjusting the parameters of the feature encoding sub-model based on the feature representation loss value, and the set conditions include that the number of times of feature encoding in the current round of clustering reaches the set number threshold, or the feature encoding sub-model meets the convergence conditions; Then clustering the multiple feature representation vectors includes: Clustering based on the multiple feature representation vectors obtained by encoding using the feature encoding sub-model after the last iterative update in the current round of clustering.
12. The method according to any one of claims 1-8, characterized in that, After outputting the multiple cluster center vectors included in the clustering result obtained in the last round of clustering, the method further includes: For each video in the video database, respectively performing the following operations: Based on the multiple cluster center vectors, respectively determining the image categories to which each image frame included in one video belongs, and associating the one video with each image category according to the determined image categories; Receiving a video retrieval request sent by a target user, where the video retrieval request carries a target video; Based on the multiple cluster center vectors, determining the image categories to which each image frame included in the target video belongs, and determining the videos associated with the determined image categories as candidate matching videos; Selecting a target matching video recommended to the target user from the candidate matching videos.
13. A clustering center determination device, characterized in that, The device includes: A sample acquisition unit, configured to acquire an image sample set and construct multiple image sample groups based on the image sample set, where each image sample group includes at least two image samples; A clustering unit, configured to perform multiple rounds of clustering using an image clustering model until the preset convergence conditions are met, where the clustering unit includes a feature encoding sub-unit, a clustering sub-unit, and a model adjustment sub-unit: The feature encoding sub-unit is configured to respectively perform feature encoding on each image sample included in the multiple image sample groups, obtain corresponding multiple feature representation vectors based on the encoding results; determine the feature representation loss value according to the similarity between at least two image samples respectively included in the multiple image sample groups; The clustering sub-unit is configured to cluster the multiple feature representation vectors and determine the clustering loss value based on the clustering result; The model adjustment sub-unit is configured to adjust the parameters of the image clustering model when it is determined that the image clustering model has not converged based on the feature representation loss value and the clustering loss value; an output unit, configured to output the multiple cluster center vectors included in the clustering result obtained in the last round of clustering, where one cluster center vector corresponds to one image category.
14. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer storage medium, on which computer program instructions are stored, characterized in that when the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Face feature extraction training method and system for face recognition
CN110490027A
Depth enhanced image clustering method
CN112464005A