Image Retrieval Method, Apparatus, Computer-Readable Medium, and Electronic Device
By performing feature extraction and semantic correlation classification prediction on images, determining target clustering clusters and performing feature comparisons, the problems of low image retrieval efficiency and poor accuracy in the prior art are solved, and efficient and accurate image retrieval effects are achieved.
Patent Information
- Application Number
- CN202110581014.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-05-26
AI Technical Summary
In the prior art, image retrieval efficiency and poor retrieval accuracy are inefficient, especially in large-scale image databases, resulting in a decrease in the breadth and accuracy of the search results.
By performing feature extraction on the query image, semantic correlation classification prediction is performed, target semantic semantic categories and target cluster clusters are determined, and feature comparison is performed to recall the target image matching the query image.
The scope of image retrieval is narrowed through secondary clustering, improve the accuracy and efficiency of image retrieval, and enable large-scale image retrieval with limited computing resources.
Smart Images

Figure CN113761261B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an image retrieval method, an image retrieval device, a computer-readable medium, and an electronic device. Background Art
[0002] Image retrieval is a commonly used technology in video and image applications. Its goal is to find other images with the same or similar content in the database based on the image content of the query image. For example, it can be used for image source identification, image recommendation or video recommendation, etc.
[0003] In order to improve the breadth of data retrieval, the database used for data retrieval generally contains a large number of image data samples, and the number of samples will continue to grow over time. Too much retrieval data will lead to a decrease in retrieval efficiency and retrieval accuracy. Summary of the invention
[0004] The purpose of this application is to provide an image retrieval method, an image retrieval device, a computer-readable medium and an electronic device, which can overcome the technical problems existing in the related art such as low retrieval efficiency and poor retrieval accuracy.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of an embodiment of the present application, an image retrieval method is provided, the method comprising: performing feature extraction on a query image to be retrieved to obtain image features of the query image; performing classification prediction on the image features to determine a target semantic category and a target clustering cluster having semantic relevance to the image features, the target clustering cluster being selected from one or more candidate clustering clusters belonging to the target semantic category; performing feature comparison between the image features and candidate images in the target clustering cluster to determine a target image matching the query image.
[0007] According to one aspect of an embodiment of the present application, an image retrieval device is provided, which includes: a feature extraction module, configured to perform feature extraction on a query image to be retrieved, and obtain image features of the query image; a classification prediction module, configured to perform classification prediction on the image features to determine a target semantic category and a target cluster cluster having semantic relevance to the image features, wherein the target cluster cluster is selected from one or more candidate cluster clusters belonging to the target semantic category; and a feature comparison module, configured to perform feature comparison on the image features and candidate images in the target cluster cluster to determine a target image that matches the query image.
[0008] In some embodiments of the present application, based on the above technical solution, the classification prediction module is configured to: obtain candidate semantic categories and candidate clustering clusters obtained by classification prediction of candidate images in an image retrieval database; perform feature comparison between the image features and each of the candidate semantic categories and candidate clustering clusters respectively to determine the target semantic categories and target clustering clusters that have semantic relevance to the image features.
[0009] In some embodiments of the present application, based on the above technical solution, the classification prediction module is also configured to: perform feature comparison between the image feature and the classification center vector of each candidate semantic category to determine a target semantic category having semantic relevance to the image feature; perform feature comparison between the image feature and the cluster center vector of each candidate cluster cluster belonging to the target semantic category to determine a target cluster cluster having semantic relevance to the image feature.
[0010] In some embodiments of the present application, based on the above technical solution, the classification prediction module is also configured to: perform feature comparison between the image feature and the classification center vector of each candidate semantic category to obtain the classification similarity between the image feature and the classification center vector; and take one or more candidate semantic categories whose classification similarity is greater than a preset similarity threshold as the target semantic category having semantic relevance to the image feature.
[0011] In some embodiments of the present application, based on the above technical solution, the classification prediction module is also configured to: perform feature splicing processing on each of the candidate semantic categories and each candidate cluster belonging to the candidate semantic category, and obtain a candidate splicing vector composed of the classification center vector of the candidate semantic category and the cluster center vector of the candidate cluster; perform feature splicing processing on the image feature with itself, and obtain a spliced image feature composed of two image features; perform feature comparison on the spliced image feature and each of the candidate splicing vectors to determine a target splicing vector that has semantic relevance to the image feature, and obtain the target semantic category and target cluster constituting the target splicing vector.
[0012] In some embodiments of the present application, based on the above technical solution, the classification prediction module is configured to: input the image features into the classification model and clustering model obtained by joint training respectively; predict the category distribution probability of the image features in multiple candidate semantic categories through the classification model, and select a target semantic category that has semantic relevance to the image features from the multiple candidate semantic categories according to the category distribution probability; predict the clustering cluster distribution probability of the image features in multiple candidate clustering clusters through the clustering model, and select a target clustering cluster that has semantic relevance to the image features from one or more candidate clustering clusters belonging to the target semantic category according to the clustering cluster distribution probability.
[0013] In some embodiments of the present application, based on the above technical solution, the classification prediction module is configured to: obtain a feature extraction model for extracting features of a query image to be retrieved and a semantic classification model and a clustering model for classifying and predicting the query image; initialize the feature extraction model, the semantic classification model and the clustering model according to preset model parameters respectively; and jointly train the feature extraction model, the semantic classification model and the clustering model based on image samples with semantic category labels to update the model parameters of each model.
[0014] In some embodiments of the present application, based on the above technical solution, the classification prediction module is configured to: perform feature extraction on image samples with semantic category labels through the feature extraction model to obtain sample features of the image samples; perform clustering processing on image samples with the same semantic category labels to obtain clustering labels of the image samples; alternately execute classification training rounds and clustering training rounds with a specified number of rounds; in the classification training rounds, jointly train the feature extraction model and the semantic classification model based on the sample features and the semantic category labels; in the clustering training rounds, jointly train the feature extraction model, the semantic classification model and the clustering model based on the sample features, the semantic category labels and the clustering labels.
[0015] In some embodiments of the present application, based on the above technical solution, the feature extraction model and the semantic classification model are jointly trained based on the sample features and the semantic category labels, including: classifying and predicting the sample features through the semantic classification model to obtain the semantic category prediction result of the image sample; determining the classification prediction error of the semantic classification model according to the semantic category label and the semantic category prediction result, and updating the model parameters of the feature extraction model and the semantic classification model according to the classification prediction error; jointly training the feature extraction model and the semantic classification model based on the sample features, the semantic category labels and the clustering labels. The semantic classification model and the clustering model include: performing classification prediction on the sample features through the semantic classification model to obtain the semantic category prediction result of the image sample; determining the classification prediction error of the semantic classification model according to the semantic category label and the semantic category prediction result; performing clustering prediction on the sample features through the clustering model to obtain the clustering prediction result of the image sample; determining the clustering prediction error of the clustering model according to the clustering label and the clustering prediction result; and updating the model parameters of the feature extraction model, the semantic classification model and the clustering model according to the classification prediction error and the clustering prediction error.
[0016] In some embodiments of the present application, based on the above technical solution, the classification prediction module is configured to: obtain one or more cluster center vectors obtained by clustering image samples with the same semantic category label in the current clustering round; obtain a cluster label sequence used as a clustering target in the previous clustering round from the clustering model; sort the cluster center vector according to the vector similarity between the cluster center vector and each cluster label in the cluster label sequence to obtain a vector sequence; and update the cluster label sequence in the clustering model according to the vector sequence.
[0017] In some embodiments of the present application, based on the above technical solution, the feature comparison module is configured to: obtain the candidate feature vector of the candidate image obtained by performing feature extraction on each candidate image in the target cluster; perform feature comparison between the image feature and the candidate feature vector to obtain the feature similarity between the image feature and the candidate feature vector; and select a target image that matches the query image from the target cluster based on the feature similarity.
[0018] In some embodiments of the present application, based on the above technical solution, the feature comparison module is also configured to: obtain a similarity threshold corresponding to the target semantic category; wherein different target semantic categories correspond to different similarity thresholds; and select a candidate image whose feature similarity is greater than the similarity threshold from the target cluster as a target image that matches the query image.
[0019] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the image retrieval method in the above technical solution is implemented.
[0020] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the image retrieval method in the above technical solution by executing the executable instructions.
[0021] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the image retrieval method in the above technical solution.
[0022] In the technical solution provided in the embodiment of the present application, by extracting features from the query image to obtain image features, the image features can be classified and predicted based on semantic relevance to obtain corresponding target semantic categories and target clusters, and further image retrieval and recall can be performed from the target clusters belonging to the target semantic category based on the image features. Image classification and clustering processing based on the secondary clustering method can narrow the image retrieval scope and improve the image retrieval accuracy.
[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 The exemplary system architecture block diagram applying the technical solution of the present application is schematically shown.
[0026] Figure 2 A block diagram schematically illustrates the principle of performing secondary clustering on an image database in an application scenario according to an embodiment of the present application.
[0027] Figure 3The principle block diagram of performing image retrieval based on a two-level clustering method in an application scenario in an embodiment of the present application is schematically shown.
[0028] Figure 4 The following is a flowchart schematically showing the steps of an image retrieval method in one embodiment of the present application.
[0029] Figure 5 A schematic diagram of the model structure composition of a feature extraction model used in one embodiment of the present application is shown.
[0030] Figure 6 A flowchart schematically shows the steps of a method for training an image processing model in one embodiment of the present application.
[0031] Figure 7 A schematic diagram of the model structure composition of the semantic classification model used in one embodiment of the present application is shown.
[0032] Figure 8 A schematic diagram of the model structure composition of the clustering model used in one embodiment of the present application is shown.
[0033] Fig. 9 The schematic diagram schematically shows the principle of reordering cluster centers in one embodiment of the present application.
[0034] Fig.10 The structural block diagram of the image retrieval device provided in an embodiment of the present application is schematically shown.
[0035] Fig.11 The structure block diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.
[0037] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0038] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0039] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0040] Figure 1 The exemplary system architecture block diagram applying the technical solution of the present application is schematically shown.
[0041] like Figure 1 As shown, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include various electronic devices such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, and a smart car terminal. The server 130 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types that can provide a communication link between the terminal device 110 and the server 130, such as a wired communication link or a wireless communication link.
[0042] According to the implementation requirements, the system architecture in the embodiment of the present application can have any number of terminal devices, networks and servers. For example, the server 130 can be a server group composed of multiple server devices. In addition, the technical solution provided in the embodiment of the present application can be applied to the terminal device 110, can also be applied to the server 130, or can be implemented by the terminal device 110 and the server 130 together, and the present application does not make any special restrictions on this.
[0043] For example, the user can upload a query image through an image retrieval client or search engine installed on the terminal device 110, thereby actively initiating an image retrieval request. After receiving the query image, the server 130 can perform an image search in the database, find other images with the same or similar content as the query image, and then return the search results to the user based on the retrieved images.
[0044] For another example, a user can watch a video through a video client or browser installed on the terminal device 110. During the video playback process of the terminal device 110, a portion of images can be extracted from the currently played video content or the historical playback record as a query image, and the query image can be uploaded to the server 130. The server 130 performs image retrieval in the database according to the query image, and after finding other images with the same or similar content as the query image, it can obtain related videos corresponding to the other images, and form a video recommendation list with the related videos, and return it to the terminal device 110, thereby realizing video content recommendation to the user.
[0045] In some embodiments of the present application, a machine learning model for image retrieval based on artificial intelligence technology can be installed on the terminal device 110 or the server 130.
[0046] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0047] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0048] Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to the use of cameras and computers to replace human eyes to identify, track and measure targets, and further perform graphics processing so that the computer processes the images into images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0049] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0050] In some embodiments of the present application, an image database for image retrieval may be configured on the terminal device 110 or the server 130. The image database may be stored on a blockchain maintained by a blockchain network. For example, the terminal device 110 or the server 130 may serve as a blockchain node constituting a blockchain network.
[0051] Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0052] The underlying blockchain platform can include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and the maintenance of the correspondence between the user's real identity and the blockchain address (authority management), etc., and, under authorization, supervises and audits the transactions of certain real identities and provides risk control rule configuration (risk control audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and records valid requests to storage after consensus is reached. For a new business request, the basic service first adapts the interface for parsing and authentication (interface adaptation), and then encrypts the business information through the consensus algorithm (consensus management). The smart contract module is responsible for the registration and issuance of contracts, as well as contract triggering and contract execution. Developers can define the contract logic in a programming language and publish it to the blockchain (contract registration). According to the logic of the contract terms, the key or other events are called to trigger the execution and complete the contract logic. It also provides the function of contract upgrade and cancellation. The operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation and real-time status visualization output of the product during the product release process, such as alarm, network monitoring, node equipment health monitoring, etc.
[0053] The platform product service layer provides the basic capabilities and implementation framework of typical applications. Developers can build on these basic capabilities and superimpose business features to complete the blockchain implementation of business logic. The application service layer provides application services based on blockchain solutions for business participants to use.
[0054] In the related technology of the present application, large-scale image retrieval can be performed using bucket retrieval. Bucket retrieval mainly divides the original large amount of data into multiple non-overlapping data sets, each data set belongs to a bucket, and during retrieval, matching image samples are found from the bucket that best matches the query image.
[0055] In some optional embodiments, the bucketing method can be generated in a clustering manner. For example, for 1 million image samples, if they are clustered into 10,000 data buckets, the cluster center is 10,000. The effect of bucketing has a great impact on the final result of the retrieval. The best bucketing result is to hope that samples with similar characteristics can be classified into the same bucket, so that the recall of a certain bucket is similar to the real sample. However, in the real data distribution, common image categories (called head categories) will appear repeatedly, while the data volume of other image categories is long-tailed, that is, the magnitude may be 1 / 100 of the head category or even less. Directly performing a global clustering bucket retrieval cannot properly handle the clustering results with a long-tail distribution, so that the images of the long-tail category and the image samples of the head category are mixed in the same clustering bucket, resulting in poor overall retrieval effect. In addition, the similarity thresholds of different types of images are actually different. For example, for scenes of fierce fighting, the features in the cluster are relatively scattered because the images change too quickly, while the face images used for face recognition are often concentrated and compact. If clustering and retrieval are performed according to a unified standard, it will be difficult to accurately recall them. Therefore, the global clustering method has certain shortcomings in terms of long-tail category clustering and image recall with different distribution densities.
[0056] The embodiment of the present application proposes a two-level clustering scheme based on semantic categories to address the problems of poor long-tail processing and difficulty in solving different distribution densities in global bucket retrieval. Through deep learning, all samples are pre-classified at the first level, and then clustering is performed within each category to obtain a second-level cluster bucket. The features of different categories are made consistent in clustering and distribution density through end-to-end classification and clustering learning. In a specific image retrieval application, the retrieval of the second level bucket can be achieved by combining classification buckets and cluster buckets. The problem of long-tail category clustering not being guaranteed is solved by clustering by category, and cluster clusters of different distribution densities are constrained to corresponding categories through the learning of the end-to-end model, so that the customized retrieval threshold can be adjusted by category according to the image category in the retrieval, solving the difficult samples of global retrieval from a finer level. The long-tail problem is controlled by the method of classification first and then clustering, and the overall retrieval time is equivalent to the original retrieval time by designing end-to-end features to support the joint retrieval of classification buckets and cluster buckets.
[0057] Figure 2 The schematic diagram shows a principle block diagram of performing secondary clustering on an image database in an application scenario in an embodiment of the present application. Figure 2As shown, for each image sample in the image database 210, its corresponding image feature 220 can be obtained through semantic inference. After deep learning classification processing of the image sample according to the image feature 220, multiple semantic categories 230 can be obtained, including category 1, category 2...category N, and each semantic category can include a corresponding number of image samples. Clustering processing is performed on the image samples belonging to the same semantic category 230 to obtain one or more cluster clusters 240 corresponding to each semantic category 230. Based on the secondary clustering, multiple semantic categories 230 and cluster clusters 240 corresponding to each semantic category 230 can be obtained. When performing image retrieval, the query range of the query image can be narrowed to the cluster cluster 240, thereby improving the image retrieval efficiency and retrieval accuracy.
[0058] Figure 3 The principle block diagram of performing image retrieval based on a two-level clustering method in an application scenario in an embodiment of the present application is schematically shown.
[0059] Based on the two-level clustering method, firstly, image features are extracted from each image sample in the image database 310, and the image features can be processed by dimensionality reduction to obtain a one-dimensional sample feature vector 320. Based on the sample feature vector 320, the image samples can be classified to obtain multiple classification centers 330 corresponding to different semantic categories, and the image samples corresponding to the same semantic category can be further clustered to obtain multiple cluster centers 340 corresponding to different cluster clusters. In the embodiment of the present application, one classification center can be associated with one or more cluster centers.
[0060] For the query image 350 to be retrieved, the embodiment of the present application uses the same feature extraction method as the image sample to extract features, and the corresponding image feature vector 360 can be obtained. The image feature vector 360 is classified and compared with each classification center 330 respectively to select the target semantic category i that matches the image feature vector 360, for example, the target semantic category is the Nth semantic category corresponding to the classification center N. The image samples under this semantic category are clustered to form three clusters, corresponding to cluster center M-2, cluster center M-1 and cluster center M respectively. The image feature vector 360 is further clustered and compared with the cluster centers corresponding to each cluster, so as to obtain the target cluster j that matches it. Subsequently, retrieval and recall can be further performed from the image samples belonging to the target semantic category i and the target cluster j to obtain the target image that matches the query image 350 features as the retrieval result.
[0061] The following is a detailed description of the technical solutions such as the image retrieval method, image retrieval device, computer-readable medium, and electronic device provided by the present application in conjunction with specific implementation methods.
[0062] Figure 4 The flowchart schematically shows the steps of the image retrieval method in one embodiment of the present application. The image retrieval method can be executed by a terminal device or a server, or by both the terminal device and the server. Figure 4 As shown, the image retrieval method in the embodiment of the present application may mainly include the following steps S410 to S430.
[0063] Step S410: extracting features of the query image to be retrieved to obtain image features of the query image.
[0064] Step S420: performing classification prediction on the image features to determine a target semantic category and a target cluster having semantic relevance to the image features, wherein the target cluster is selected from one or more candidate clusters belonging to the target semantic category.
[0065] Step S430: performing feature comparison between the image features and the candidate images in the target cluster to determine a target image that matches the query image.
[0066] In the image retrieval method provided in the embodiment of the present application, by extracting features from the query image to obtain image features, the image features can be classified and predicted based on semantic relevance to obtain corresponding target semantic categories and target clusters, and further image retrieval and recall can be performed from the target clusters belonging to the target semantic category based on the image features. Image classification and clustering processing based on the secondary clustering method can narrow the image retrieval scope and improve the image retrieval accuracy.
[0067] The specific implementation of each method step in the image retrieval method is described in detail below.
[0068] In step S410, feature extraction is performed on the query image to be retrieved to obtain image features of the query image.
[0069] In one embodiment of the present application, a method for extracting features from a query image may be to perform data mapping on the query image through a pre-trained feature extraction model to obtain an embedding vector embedding.
[0070] Figure 5 FIG. 2 shows a schematic diagram of the model structure composition of a feature extraction model used in one embodiment of the present application. Figure 5As shown in the figure, the feature extraction model is a neural network model built on the basis of ResNet-101. The model mainly includes five convolutional layers, namely Conv1, Conv2_x, Conv3_x, Conv4_x, and Conv5_x, which are connected from the input end to the output end in sequence. The last convolutional layer Conv5_x can output the embedding vector embedding as the image feature. Among them, the convolutional layer Conv1 uses a convolutional kernel with a size of 7X7 and a number of channels of 64 to convolve the input image with a step size of 2 to obtain a feature map with an output size of 300X500. The convolutional layer Conv2_x first uses a pooling window with a size of 3X3 to perform maximum pooling processing on the feature map output by Conv1 with a step size of 2, and then convolutions through three identical convolutional blocks in sequence. Among them, each convolutional block includes convolutional kernels of sizes 1X1, 3X3, and 1X1 in sequence to convolve the input feature map. The subsequent convolutional layers Conv3_x, Conv4_x, and Conv5_x have similar convolutional block structures and will not be described here.
[0071] In one embodiment of the present application, after the last convolution layer Conv5_x, the output data can be pooled through a pooling layer, and the output data after pooling can be further vector normalized through a normalization layer, so as to facilitate subsequent feature comparison with the classification center and the vector center.
[0072] In step S420, classification prediction is performed on the image features to determine a target semantic category and a target cluster having semantic relevance to the image features. The target cluster is selected from one or more candidate clusters belonging to the target semantic category.
[0073] In one embodiment of the present application, the method for classifying and predicting image features may be to perform semantic matching detection on the image features with each semantic category and each image clustering cluster. The embodiment of the present application may obtain candidate semantic categories and candidate clustering clusters obtained by classifying and predicting candidate images in an image retrieval database, and then perform feature comparison on the image features with each candidate semantic category and candidate clustering cluster respectively to determine the target semantic category and target clustering cluster that have semantic relevance to the image features. The process and principle of classifying and predicting candidate images in an image retrieval database may be referred to Figure 2 , I will not go into details here.
[0074] In one embodiment of the present application, a classification center vector may be determined for each candidate language category, and a cluster center vector may be determined for each candidate cluster. On this basis, the method of comparing the image features with each candidate semantic category and candidate cluster may be to perform similarity matching between the image features and each classification center vector and cluster center vector. Similarity matching may be performed on the image features using second-order matching or first-order matching.
[0075] In one embodiment of the present application, the quantity ratio between the candidate clusters and the candidate semantic categories can be obtained. When the quantity ratio is greater than a preset ratio threshold, a second-order matching method can be used for similarity matching; when the quantity ratio is less than or equal to the preset ratio threshold, a first-order matching method can be used for similarity matching. The ratio threshold can be, for example, 2.
[0076] In an embodiment of similarity matching using second-order matching, a method for comparing image features with each candidate semantic category and candidate clustering cluster may include: firstly comparing image features with classification center vectors of each candidate semantic category to determine a target semantic category having semantic relevance to the image features; and then comparing image features with cluster center vectors of each candidate clustering cluster belonging to the target semantic category to determine a target clustering cluster having semantic relevance to the image features. When the number of candidate clustering clusters is large, for example, exceeding a specified multiple of the number of candidate semantic categories, using second-order matching can gradually narrow the matching range, reduce the amount of data calculation, and thereby improve the data processing efficiency of feature comparison.
[0077] In one embodiment of the present application, a method for performing feature comparison between an image feature and a classification center vector of each candidate semantic category may include: performing feature comparison between an image feature and a classification center vector of each candidate semantic category to obtain a classification similarity between the image feature and the classification center vector; and taking one or more candidate semantic categories whose classification similarity is greater than a preset similarity threshold as a target semantic category having semantic relevance to the image feature. When the classification similarities between multiple classification center vectors and image features are all greater than the preset similarity threshold, it indicates that the image feature has a high semantic similarity with multiple candidate semantic categories, and at this time, multiple candidate semantic categories can all be selected as target semantic categories. Since the scope can be further narrowed down by cluster matching in the future, although selecting multiple candidate semantic categories as target semantic categories may introduce some irrelevant samples, it will not affect the final retrieval accuracy.
[0078] In an embodiment of similarity matching using first-order matching, the method of comparing image features with each candidate semantic category and candidate cluster cluster may include: performing feature splicing processing on each candidate semantic category and each candidate cluster cluster belonging to the candidate semantic category, obtaining a candidate splicing vector composed of a classification center vector of the candidate semantic category and a cluster center vector of the candidate cluster cluster; performing feature splicing processing on the image feature with itself, obtaining a spliced image feature composed of two image features; performing feature comparison on the spliced image feature with each candidate splicing vector to determine a target splicing vector having semantic relevance to the image feature, and obtaining a target semantic category and a target cluster cluster constituting the target splicing vector. When the number of candidate cluster clusters is relatively small, for example, less than a specified multiple of the number of candidate semantic categories, using first-order matching can reduce the amount of data calculation for vector calculation, thereby improving the data processing efficiency of feature comparison.
[0079] In one embodiment of the present application, the method for classifying and predicting image features may include performing model prediction using a pre-trained image processing model. The image processing model may include a feature extraction model for extracting image features, a classification model for semantically classifying images, and a clustering model for clustering images with the same semantics.
[0080] In one embodiment of the present application, the method for classifying and predicting image features in step S420 may include: respectively inputting the image features into a classification model and a clustering model obtained by joint training; predicting the category distribution probability of the image features in multiple candidate semantic categories through the classification model, and selecting a target semantic category having semantic relevance to the image features from the multiple candidate semantic categories based on the category distribution probability; predicting the clustering cluster distribution probability of the image features in multiple candidate clustering clusters through the clustering model, and selecting a target clustering cluster having semantic relevance to the image features from one or more candidate clustering clusters belonging to the target semantic category based on the clustering cluster distribution probability.
[0081] Figure 6 The following is a schematic flow chart showing the steps of a method for training an image processing model in one embodiment of the present application. Figure 6 As shown, in an embodiment of the present application, the method for training an image processing model may include the following steps S610 to S630.
[0082] Step S610: Acquire a feature extraction model for extracting features from a query image to be retrieved and a semantic classification model and a clustering model for classifying and predicting the query image.
[0083] The feature extraction model can be Figure 5 The ResNet-101 based neural network model shown in .
[0084] Figure 7 FIG. 2 shows a schematic diagram of the structure of a semantic classification model used in one embodiment of the present application. Figure 7 As shown, the semantic classification model may include a pooling layer Pool_cr, a normalization layer Norm, and a fully connected layer Fc_cr connected in sequence. Wherein, N represents the number of semantic categories to be learned, which is related to the image data type and the number of image samples in the actual application scenario. N can be a positive integer, for example, N can be a specified value between 100 and 500. The network parameter of the fully connected layer Fc_cr is a weight matrix of 2048*N, corresponding to N classification center vectors with a vector length of 2048.
[0085] Figure 8 The schematic diagram of the model structure composition of the clustering model used in one embodiment of the present application is shown. The clustering model can be used with Figure 7 The model structure is similar to the classification model shown in FIG. 1 , that is, it includes a pooling layer Pool_cluster, a normalization layer Norm, and a fully connected layer Fc_cluster connected in sequence. Among them, M represents the number of semantic categories to be learned, which is related to the image data type and the number of image samples in the actual application scenario. M can be a positive integer, for example, M can be a specified value between 500 and 1000. The network parameter of the fully connected layer Fc_cluster is a weight matrix of 2048*M, corresponding to M cluster center vectors with a vector length of 2048.
[0086] Step S620: Initialize the feature extraction model, semantic classification model and clustering model according to preset model parameters respectively.
[0087] In one embodiment of the present application, the feature extraction model can be initialized with pre-trained model parameters, such as Imagenet pre-trained classification parameters, or parameters obtained by training retrieval features, etc. The fully connected layer in the semantic classification model and clustering model can be initialized with parameters that conform to a Gaussian distribution with a preset variance and a preset mean. The preset variance can be, for example, 0.01, and the preset mean can be, for example, 0.
[0088] Step S630: jointly train the feature extraction model, the semantic classification model and the clustering model based on the image samples with semantic category labels to update the model parameters of each model.
[0089] In one embodiment of the present application, a method for jointly training a feature extraction model, a semantic classification model, and a clustering model based on image samples with semantic category labels may include: performing feature extraction on image samples with semantic category labels through a feature extraction model to obtain sample features of the image samples; performing clustering processing on image samples with the same semantic category labels to obtain clustering labels of the image samples; alternatingly executing classification training rounds and clustering training rounds with a specified number of rounds; in the classification training rounds, jointly training the feature extraction model and the semantic classification model based on sample features and semantic category labels; in the clustering training rounds, jointly training the feature extraction model, the semantic classification model, and the clustering model based on sample features, semantic category labels, and clustering labels.
[0090] In one embodiment of the present application, a method for jointly training a feature extraction model and a semantic classification model based on sample features and semantic category labels may include: performing classification prediction on sample features through a semantic classification model to obtain a semantic category prediction result of the image sample; determining the classification prediction error of the semantic classification model based on the semantic category label and the semantic category prediction result, and updating the model parameters of the feature extraction model and the semantic classification model based on the classification prediction error.
[0091] In the embodiment of the present application, all parameters of the feature extraction model and the semantic classification model are set to a state requiring learning. During training, the neural network performs forward calculation on an input image to obtain a classification prediction result, and compares it with the annotated category label to calculate the classification loss value of the model. The classification loss loss performs a gradient backward calculation to obtain the updated values of all model parameters, and updates the model parameters of the feature extraction model and the semantic classification model.
[0092] In one embodiment of the present application, a method for jointly training a feature extraction model, a semantic classification model, and a clustering model based on sample features, semantic category labels, and clustering labels may include: performing classification prediction on sample features through a semantic classification model to obtain a semantic category prediction result of an image sample; determining a classification prediction error of the semantic classification model based on the semantic category label and the semantic category prediction result; performing clustering prediction on sample features through a clustering model to obtain a clustering prediction result of the image sample; determining a clustering prediction error of the clustering model based on the clustering label and the clustering prediction result; and updating model parameters of the feature extraction model, the semantic classification model, and the clustering model based on the classification prediction error and the clustering prediction error.
[0093] In an embodiment of the present application, for category i, for example, kmeans clustering can be performed using samples of the category to obtain Mi cluster centers, and the number of clusters is Mi (the number of cluster centers for different categories is different). After obtaining the cluster centers of all categories, the corresponding cluster center labels are assigned to all samples. All parameters of the feature extraction model, semantic classification model, and clustering model are set to a learning state. During training, the neural network performs a forward calculation on an input image to obtain the classification prediction result and the clustering prediction result, and compares the total classification loss value (classification loss) of the model with the annotated category label and cluster category. The total classification loss loss performs a gradient backward calculation to obtain the updated values of all model parameters, and updates the model parameters of the feature extraction model, semantic classification model, and clustering model.
[0094] In one embodiment of the present application, after re-clustering the image samples, it is necessary to synchronously update the relevant parameters in the clustering model according to the clustering results. In the embodiment of the present application, the method for updating the relevant parameters of the clustering model based on the re-clustering results may include: obtaining one or more cluster center vectors obtained by clustering the image samples with the same semantic category label in the current clustering round; obtaining the cluster label sequence used as the clustering target in the previous clustering round from the clustering model; sorting the cluster center vectors according to the vector similarity between the cluster center vectors and each cluster label in the cluster label sequence to obtain a vector sequence; and updating the cluster label sequence in the clustering model according to the vector sequence.
[0095] In one embodiment of the present application, when clustering image samples, the number of clusters may be determined first, and then the cluster centers that meet the number of clusters may be determined.
[0096] For samples in each category, the number of cluster centers needs to be determined in advance. Assuming that M categories need to be clustered globally, the number of all samples is Q, Q can be a positive integer, and assuming that the number of samples in the i-th category is Si, then the number of centers required for clustering in this category is as follows: Ci, that is, the number of cluster centers obtained according to the proportion of the data volume, at least one cluster center is required (even for categories with very little data).
[0097] Ci=Max(Si / (Q / M),1)
[0098] For the image samples in each category, kmeans clustering can be performed according to the cluster centers determined above to obtain the corresponding number of cluster centers; after completing the clustering of all categories, the cluster numbers of each cluster center are recorded in the order of categories (such as Figure 2 The clusters 1, 2, …, M in the cluster are obtained to obtain the initial cluster center Ncluster of this round, with a total of M*2048 cluster vectors.
[0099] When clustering is performed for the first time in the first round of model iteration, the fully connected layer Fc_cluster (M*2048) in the clustering model is initialized with a Gaussian distribution. At this time, the cluster centers in the above steps can be copied to Fc_cluster; at the same time, the cluster center numbers corresponding to each classification category are recorded (such as Figure 2 In the above example, category 1 corresponds to clusters 1 and 2, and category N corresponds to clusters M-2, M-1, and M).
[0100] If it is not the first clustering, Ncluster is reordered and copied to Fc_cluster. At this time, Fc_cluster records the center of the last clustering, and each center of the new Ncluster can be reordered based on the principle of selecting the closest center based on cosine similarity and the M centers of the last clustering in Fc_cluster. Fig. 9 The schematic diagram schematically shows the principle of reordering cluster centers in one embodiment of the present application. Fig. 9 As shown in the figure, the cluster centers of the previous round form the sequence P1, P2, P3, and P4. After re-clustering, the new cluster centers are Pnew1, Pnew2, Pnew3, and Pnew4. For example, if P1 and Pnew4 are the most similar centers, then after Ncluster is re-sorted, Pnew4 is the first, and so on to complete the new re-sorting in other clusters. After re-sorting, copy the new cluster center Ncluster2 to Fc_cluster, and record the cluster center number corresponding to each classification category.
[0101] According to the continuously updated cluster centers, for each training sample, find the nearest cluster center i in the category to which it belongs, and assign the sample the cluster label corresponding to the cluster center i (the serial number of the center, between 1...M).
[0102] The purpose of reordering for non-first clustering is to maintain the maximum similarity of the previous and next cluster IDs. The more similar the previous and next clusters are, the more stable the cluster category to which the sample belongs is. Stable cluster categories are very important for deep learning convergence. If reordering is not performed, the cluster category to which the sample belongs will change greatly after each re-clustering, and the loss will fluctuate greatly. This will interfere with the embedding trained last time and cause the embedding to have to start learning again for a long time to reach the last convergence level.
[0103] The model training process of an embodiment of the present application in an application scenario is as follows.
[0104] 1) Classification learning is performed first. After completing round E (such as the 10th round) of classification learning, the feature extraction model and the semantic classification model are relatively stable.
[0105] 2) In the E+1 round, the model is trained by the joint training method of classification learning and clustering learning, that is, the weighted sum of the classification loss and clustering loss of the model is calculated as the final loss. In fact, the focus of this round is to make the embedding cluster-friendly (with a more reasonable cluster distribution) through clustering. However, in order to avoid excessive loss fluctuations due to changes in cluster centers, it is necessary to add classification loss control at the same time, L = (1-a)Lclass + a*Lcluster, where a can be selected according to the situation, for example, 0.8.
[0106] 4) In the E+2 round, classification learning is performed.
[0107] 5) In the E+3 round, classification + clustering learning is performed.
[0108] …
[0109] After rounds of alternating iterative training, the model convergence condition is finally reached. The purpose of alternating classification and clustering training during this training process is to protect the results of classification and avoid excessive fluctuations in classification and embedding due to fluctuations in clustering.
[0110] In step S430, the image features are compared with the candidate images in the target cluster to determine the target image that matches the query image.
[0111] In one embodiment of the present application, a method for performing feature comparison between image features and candidate images in a target cluster may include: obtaining a candidate feature vector of the candidate image obtained by performing feature extraction on each candidate image in the target cluster; performing feature comparison between the image features and the candidate feature vector to obtain feature similarity between the image features and the candidate feature vector; and selecting a target image that matches the query image from the target cluster based on the feature similarity.
[0112] In one embodiment of the present application, a method for selecting a target image that matches a query image from a target cluster based on feature similarity may include: obtaining a similarity threshold corresponding to a target semantic category; wherein different target semantic categories correspond to different similarity thresholds; and selecting a candidate image whose feature similarity is greater than the similarity threshold from the target cluster as a target image that matches the query image.
[0113] In the embodiment of the present application, different similarity thresholds can be configured for different semantic categories, so that images of semantic categories with different distribution characteristics can achieve customized retrieval threshold adjustment.
[0114] The embodiments of the present application are also applicable to clustering with limited computing resources, that is, by decomposing the full amount of data into different categories, the number of clustering samples required each time is reduced, thereby allowing us to perform clustering in a limited computer memory space. Assuming that the data volume of global image samples is 100 million data, a query image has a high degree of match with 5 semantic categories out of 1,000 semantic categories, and assuming that the number of image samples of these 5 categories is 10,000, then the overall sample size of image retrieval can be reduced by 4 orders of magnitude. Therefore, the embodiments of the present application can realize large-scale clustering and image retrieval with limited computing resources.
[0115] In practical applications, the embodiments of the present application can recall stock images through two-level bucket retrieval. First, the classification center obtained after the image processing model is trained, and the cluster center corresponding to each classification center is recorded; secondly, in the registration stage, a retrieval library is established for all samples, that is, the features of the samples, the classification centers corresponding to the features, and the cluster centers corresponding to the features (searched within the classification centers) are extracted, and the stock sample id corresponding to each cluster center is recorded. Finally, in the retrieval stage, for the query graph to be retrieved, the image features of the query graph are first extracted, and the corresponding classification center i is found according to the image features, and the Ri cluster centers corresponding to the classification center i are obtained, and the cluster center m with the closest image features among the Ri cluster centers is found, and the stock samples are obtained as the result of the bucket retrieval according to the sample records of each cluster center in the registration stage.
[0116] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0117] The following introduces an apparatus embodiment of the present application, which can be used to execute the image retrieval method in the above-mentioned embodiment of the present application. Fig.10 The structure block diagram of the image retrieval device provided in the embodiment of the present application is schematically shown. Fig.10As shown, the image retrieval device 1000 includes: a feature extraction module 1010, configured to perform feature extraction on a query image to be retrieved to obtain image features of the query image; a classification prediction module 1020, configured to perform classification prediction on the image features to determine a target semantic category and a target cluster cluster having semantic relevance to the image features, wherein the target cluster cluster is selected from one or more candidate cluster clusters belonging to the target semantic category; and a feature comparison module 1030, configured to perform feature comparison on the image features and candidate images in the target cluster cluster to determine a target image that matches the query image.
[0118] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is configured to: obtain candidate semantic categories and candidate clustering clusters obtained by classification prediction of candidate images in the image retrieval database; perform feature comparison between the image features and each of the candidate semantic categories and candidate clustering clusters respectively to determine the target semantic categories and target clustering clusters that have semantic relevance to the image features.
[0119] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is further configured to: perform feature comparison between the image feature and the classification center vector of each candidate semantic category to determine a target semantic category having semantic relevance to the image feature; perform feature comparison between the image feature and the cluster center vector of each candidate cluster cluster belonging to the target semantic category to determine a target cluster cluster having semantic relevance to the image feature.
[0120] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is also configured to: perform feature comparison between the image feature and the classification center vector of each candidate semantic category to obtain the classification similarity between the image feature and the classification center vector; and take one or more candidate semantic categories whose classification similarity is greater than a preset similarity threshold as target semantic categories that have semantic relevance to the image feature.
[0121] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is further configured to: perform feature splicing processing on each of the candidate semantic categories and each of the candidate clustering clusters belonging to the candidate semantic category, and obtain a candidate splicing vector composed of the classification center vector of the candidate semantic category and the clustering center vector of the candidate clustering cluster; perform feature splicing processing on the image feature with itself, and obtain a spliced image feature composed of two image features; perform feature comparison on the spliced image feature and each of the candidate splicing vectors to determine a target splicing vector that has semantic relevance to the image feature, and obtain a target semantic category and a target clustering cluster that constitute the target splicing vector.
[0122] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is configured to: input the image features into the classification model and clustering model obtained by joint training respectively; predict the category distribution probability of the image features in multiple candidate semantic categories through the classification model, and select a target semantic category that has semantic relevance to the image features from the multiple candidate semantic categories according to the category distribution probability; predict the clustering cluster distribution probability of the image features in multiple candidate clustering clusters through the clustering model, and select a target clustering cluster that has semantic relevance to the image features from one or more candidate clustering clusters belonging to the target semantic category according to the clustering cluster distribution probability.
[0123] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is also configured to: obtain a feature extraction model for extracting features of a query image to be retrieved and a semantic classification model and a clustering model for classifying and predicting the query image; initialize the feature extraction model, the semantic classification model and the clustering model according to preset model parameters respectively; and jointly train the feature extraction model, the semantic classification model and the clustering model based on image samples with semantic category labels to update the model parameters of each model.
[0124] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is also configured to: perform feature extraction on image samples with semantic category labels through the feature extraction model to obtain sample features of the image samples; perform clustering processing on image samples with the same semantic category labels to obtain clustering labels of the image samples; alternately execute classification training rounds and clustering training rounds with a specified number of rounds; in the classification training rounds, jointly train the feature extraction model and the semantic classification model based on the sample features and the semantic category labels; in the clustering training rounds, jointly train the feature extraction model, the semantic classification model and the clustering model based on the sample features, the semantic category labels and the clustering labels.
[0125] In some embodiments of the present application, based on the above embodiments, the feature extraction model and the semantic classification model are jointly trained based on the sample features and the semantic category labels, including: performing classification prediction on the sample features through the semantic classification model to obtain the semantic category prediction result of the image sample; determining the classification prediction error of the semantic classification model according to the semantic category label and the semantic category prediction result, and updating the model parameters of the feature extraction model and the semantic classification model according to the classification prediction error; jointly training the feature extraction model, the semantic category label and the clustering label based on the sample features, the semantic category label and the clustering label. The semantic classification model and the clustering model include: performing classification prediction on the sample features through the semantic classification model to obtain the semantic category prediction result of the image sample; determining the classification prediction error of the semantic classification model according to the semantic category label and the semantic category prediction result; performing clustering prediction on the sample features through the clustering model to obtain the clustering prediction result of the image sample; determining the clustering prediction error of the clustering model according to the clustering label and the clustering prediction result; and updating the model parameters of the feature extraction model, the semantic classification model and the clustering model according to the classification prediction error and the clustering prediction error.
[0126] In some embodiments of the present application, based on the above embodiments, the classification prediction module 1020 is also configured to: obtain one or more cluster center vectors obtained by clustering image samples with the same semantic category label in the current clustering round; obtain a cluster label sequence used as a clustering target in the previous clustering round from the clustering model; sort the cluster center vector according to the vector similarity between the cluster center vector and each cluster label in the cluster label sequence to obtain a vector sequence; and update the cluster label sequence in the clustering model according to the vector sequence.
[0127] In some embodiments of the present application, based on the above embodiments, the feature comparison module 1030 is configured to: obtain a candidate feature vector of the candidate image obtained by performing feature extraction on each candidate image in the target cluster; perform feature comparison between the image feature and the candidate feature vector to obtain feature similarity between the image feature and the candidate feature vector; and select a target image that matches the query image from the target cluster based on the feature similarity.
[0128] In some embodiments of the present application, based on the above embodiments, the feature comparison module 1030 is also configured to: obtain a similarity threshold corresponding to the target semantic category; wherein different target semantic categories correspond to different similarity thresholds; and select a candidate image whose feature similarity is greater than the similarity threshold from the target cluster as a target image that matches the query image.
[0129] The specific details of the image retrieval device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments and will not be repeated here.
[0130] Fig.11 The computer system structure block diagram of the electronic device used to implement the embodiment of the present application is schematically shown. The electronic device may be as follows: Figure 1 The terminal device 110 or the server 130 shown in FIG.
[0131] It should be noted that Fig.11 The computer system 1100 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0132] like Fig.11 As shown, the computer system 1100 includes a central processing unit 1101 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1102 (ROM) or the program loaded from the storage part 1108 to the random access memory 1103 (RAM). Various programs and data required for system operation are also stored in the random access memory 1103. The central processing unit 1101, the read-only memory 1102 and the random access memory 1103 are connected to each other through a bus 1104. An input / output interface 1105 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1104.
[0133] In some embodiments, the following components are connected to the input / output interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a local area network card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed, so that a computer program read therefrom is installed into the storage section 1108 as needed.
[0134] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1109, and / or installed from the removable medium 1111. When the computer program is executed by the central processor 1101, various functions defined in the system of the present application are executed.
[0135] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0136] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0138] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable an electronic device to execute the method according to the implementation methods of the present application.
[0139] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application.
[0140] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. An image retrieval method, It is characterized in that include: Performing feature extraction on image samples with semantic category labels by using a feature extraction model to obtain sample features of the image samples; Performing clustering processing on image samples with the same semantic category labels to obtain cluster labels of the image samples; Alternately performing classification training rounds and clustering training rounds with a specified number of rounds; in the classification training rounds, jointly training the feature extraction model and the semantic classification model based on the sample features and the semantic category labels; in the clustering training rounds, jointly training the feature extraction model, the semantic classification model, and the clustering model based on the sample features, the semantic category labels, and the clustering labels; According to the feature extraction model, feature extraction is performed on the query image to be retrieved to obtain image features of the query image; According to the semantic classification model and the clustering model, the image features are classified and predicted to determine a target semantic category and a target clustering cluster having semantic relevance to the image features, wherein the target clustering cluster is selected from one or more candidate clustering clusters belonging to the target semantic category; The image features are compared with the candidate images in the target cluster to determine a target image that matches the query image.
2. The image retrieval method according to claim 1, It is characterized in that After obtaining the image features of the query image, the method further includes: Obtain candidate semantic categories and candidate clusters obtained by classifying and predicting candidate images in an image retrieval database; The image features are compared with each of the candidate semantic categories and candidate clusters to determine a target semantic category and a target cluster that have semantic relevance to the image features.
3. The image retrieval method according to claim 2, It is characterized in that The image feature is compared with each of the candidate semantic categories and the candidate clusters to determine a target semantic category and a target cluster having semantic relevance to the image feature, including: Performing feature comparison between the image feature and the classification center vector of each candidate semantic category to determine a target semantic category having semantic relevance to the image feature; The image feature is compared with the cluster center vectors of each candidate cluster belonging to the target semantic category to determine a target cluster having semantic relevance to the image feature.
4. The image retrieval method according to claim 3, It is characterized in that Performing feature comparison between the image feature and the classification center vector of each candidate semantic category to determine a target semantic category having semantic relevance to the image feature includes: Performing feature comparison between the image feature and the classification center vector of each candidate semantic category to obtain classification similarity between the image feature and the classification center vector; One or more candidate semantic categories whose classification similarity is greater than a preset similarity threshold are used as target semantic categories having semantic relevance to the image feature.
5. The image retrieval method according to claim 2, It is characterized in that The image feature is compared with each of the candidate semantic categories and the candidate clusters to determine a target semantic category and a target cluster having semantic relevance to the image feature, including: Performing feature concatenation processing on each of the candidate semantic categories and each of the candidate clusters belonging to the candidate semantic category, respectively, to obtain a candidate concatenation vector composed of the classification center vector of the candidate semantic category and the cluster center vector of the candidate cluster; Performing feature splicing processing on the image feature and itself to obtain a spliced image feature composed of the two image features; The stitched image feature is compared with each of the candidate stitching vectors to determine a target stitching vector having semantic relevance to the image feature, and a target semantic category and a target cluster constituting the target stitching vector are obtained.
6. The image retrieval method according to claim 1, It is characterized in that According to the semantic classification model and the clustering model, the image features are classified and predicted to determine a target semantic category and a target clustering cluster having semantic relevance to the image features, including: Inputting the image features into the semantic classification model and the clustering model obtained by joint training respectively; Predicting the category distribution probability of the image feature in a plurality of candidate semantic categories by using the semantic classification model, and selecting a target semantic category having semantic relevance to the image feature from the plurality of candidate semantic categories according to the category distribution probability; The clustering model is used to predict the clustering cluster distribution probability of the image feature in multiple candidate clustering clusters, and a target clustering cluster having semantic relevance to the image feature is selected from one or more candidate clustering clusters belonging to the target semantic category according to the clustering cluster distribution probability.
7. The image retrieval method according to claim 6, It is characterized in that Before extracting features from image samples with semantic category labels using a feature extraction model, the method further includes: Acquire a feature extraction model for extracting features of a query image to be retrieved and a semantic classification model and a clustering model for classifying and predicting the query image; The feature extraction model, the semantic classification model and the clustering model are initialized according to preset model parameters respectively.
8. The image retrieval method according to claim 1, It is characterized in that Jointly training the feature extraction model and the semantic classification model based on the sample features and the semantic category labels includes: Performing classification prediction on the sample features by using the semantic classification model to obtain a semantic category prediction result of the image sample; Determining a classification prediction error of the semantic classification model according to the semantic category label and the semantic category prediction result, and updating model parameters of the feature extraction model and the semantic classification model according to the classification prediction error; Jointly training the feature extraction model, the semantic classification model and the clustering model based on the sample features, the semantic category labels and the clustering labels includes: Performing classification prediction on the sample features by using the semantic classification model to obtain a semantic category prediction result of the image sample; Determining a classification prediction error of the semantic classification model according to the semantic category label and the semantic category prediction result; Performing cluster prediction on the sample features by using the cluster model to obtain a cluster prediction result of the image sample; Determining a clustering prediction error of the clustering model according to the clustering label and the clustering prediction result; Model parameters of the feature extraction model, the semantic classification model, and the clustering model are updated according to the classification prediction error and the clustering prediction error.
9. The image retrieval method according to claim 1, It is characterized in that After clustering the image samples having the same semantic category label to obtain the cluster labels of the image samples, the method further includes: Obtain one or more cluster center vectors obtained by clustering image samples with the same semantic category label in the current clustering round; Obtaining a clustering label sequence used as a clustering target in a previous clustering round from the clustering model; According to the vector similarity between the cluster center vector and each cluster label in the cluster label sequence, the cluster center vector is sorted to obtain a vector sequence; A clustering label sequence in the clustering model is updated according to the vector sequence.
10. The image retrieval method according to any one of claims 1 to 9, It is characterized in that Comparing the image features with the candidate images in the target cluster to determine a target image that matches the query image includes: Acquire a candidate feature vector of the candidate image obtained by performing feature extraction on each candidate image in the target cluster; Performing feature comparison between the image feature and the candidate feature vector to obtain feature similarity between the image feature and the candidate feature vector; A target image matching the query image is selected from the target cluster according to the feature similarity.
11. The image retrieval method according to claim 10, It is characterized in that Selecting a target image matching the query image from the target cluster according to the feature similarity includes: Acquire a similarity threshold value corresponding to the target semantic category; wherein different target semantic categories correspond to different similarity threshold values; A candidate image whose feature similarity is greater than the similarity threshold is selected from the target cluster as a target image matching the query image.
12. An image retrieval device, It is characterized in that include: A feature extraction module is configured to extract features of image samples with semantic category labels through a feature extraction model to obtain sample features of the image samples; Performing clustering processing on image samples with the same semantic category label to obtain cluster labels of the image samples; alternately performing classification training rounds and clustering training rounds with a specified number of rounds; in the classification training rounds, jointly training the feature extraction model and the semantic classification model based on the sample features and the semantic category labels; in the clustering training rounds, jointly training the feature extraction model, the semantic classification model and the clustering model based on the sample features, the semantic category labels and the cluster labels; performing feature extraction on a query image to be retrieved according to the feature extraction model to obtain image features of the query image; a classification prediction module, configured to perform classification prediction on the image feature according to the semantic classification model and the clustering model to determine a target semantic category and a target clustering cluster having semantic relevance to the image feature, wherein the target clustering cluster is selected from one or more candidate clustering clusters belonging to the target semantic category; The feature comparison module is configured to compare the image features with the candidate images in the target cluster to determine a target image that matches the query image.
13. A computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the image retrieval method according to any one of claims 1 to 11.
14. An electronic device, It is characterized in that include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to perform the image retrieval method described in any one of claims 1 to 11 by executing the executable instructions.
15. A computer program product comprising computer instructions, It is characterized in that When the computer instructions are executed by a processor, the image retrieval method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Image retrieval method and device, medium and electronic equipment
CN110297935A