Image retrieval apparatus and method, and non-volatile computer-readable storage medium
Patent Information
- Application Number
- CN202380010840.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the accuracy of image retrieval is limited by the inaccurate features extracted by neural networks such as CNN, which makes it difficult to accurately find the corresponding pages of the image in the picture book or distinguish between two people wearing the same clothes.
CNN is used for preliminary image search, and a collection of candidate images similar to the images to be retrieved are obtained. Then, further screening is performed in the image base library through reRank technology and self-attention mechanism model to expand the candidate image collection to determine the precise search results.
The accuracy of image retrieval is improved, error screening is avoided due to insufficient feature distinction, and the correct search results are included in the preliminary search results.
Smart Images

Figure CN120035848A_ABST
Abstract
Description
Image retrieval device, method and non-volatile computer-readable storage medium Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to an image retrieval device, an image retrieval method, and a non-volatile computer-readable storage medium. Background Art
[0002] Image retrieval technology can be used to retrieve images matching the probe image from multiple gallery images as search results. Image retrieval technology can be applied to a variety of scenarios. For example, when comparing images in a picture book, image retrieval technology can be used to find the page corresponding to an image on a page in the image database.
[0003] In related technologies, image retrieval can be performed based on image features extracted by neural network models such as CNN.
[0004] Summary of the Invention
[0005] According to some embodiments of the present disclosure, an image retrieval device is provided, comprising at least one processor, wherein the at least one processor is configured to: determine a set of candidate images from a plurality of comparison images based on a degree of similarity between an image to be retrieved and each of a plurality of comparison images; and determine a retrieval result of the image to be retrieved from the set of candidate images based on a degree of similarity between the plurality of comparison images.
[0006] In some embodiments, the similarity between the multiple compared images includes the similarity between each candidate image in the candidate image set and each compared image, and / or the similarity between multiple candidate images in the candidate image set.
[0007] In some embodiments, determining the retrieval result of the image to be retrieved from the candidate image set based on the similarity between multiple comparison images includes: determining at least one similar image corresponding to each candidate image from the multiple comparison images based on the similarity between each candidate image and the comparison image; expanding the candidate image set based on each candidate image and its at least one similar image; and determining the retrieval result from the expanded candidate image set.
[0008] In some embodiments, expanding the candidate image set based on each candidate image and at least one similar image includes: calculating the number of first images belonging to the candidate image set among at least one similar image corresponding to any candidate image; and determining whether to expand the at least one similar image corresponding to any candidate image to the candidate image set based on the number of first images.
[0009] In some embodiments, determining whether to expand at least one similar image corresponding to any candidate image to the candidate image set based on the number of images includes: when the proportion of the first image number in the second image number is greater than a threshold, expanding at least one similar image corresponding to any candidate image to the candidate image set, and the second image number is the number of at least one similar image corresponding to any candidate image.
[0010] In some embodiments, based on the degree of similarity between multiple comparison images, determining the retrieval results of the image to be retrieved from the candidate image set includes: using a self-attention mechanism model to process each candidate image in the candidate image set to determine the self-attention feature information of each candidate image; determining the retrieval results based on the self-attention feature information of each candidate image and the feature information of the image to be retrieved.
[0011] In some embodiments, using a self-attention mechanism model to process each candidate image in a candidate image set to determine the self-attention feature information of each candidate image includes: extracting self-attention feature information using a self-attention mechanism model based on feature vectors of multiple candidate images located in a self-attention window in a feature vector sequence, the feature vector sequence including the feature vector of each candidate image; sliding the self-attention window on the feature vector sequence according to a sliding step; repeating the above extraction step and sliding step until the feature vectors of all candidate images in the feature vector sequence are processed.
[0012] In some embodiments, the length of the self-attention window is greater than the sliding step size.
[0013] In some embodiments, determining the retrieval result based on the self-attention feature information of each candidate image and the feature information of the image to be retrieved includes: determining the cross-attention feature information of the image to be retrieved using a cross-attention model based on the feature vector of the image to be retrieved and the feature vector of each candidate image; and determining the retrieval result based on the cross-attention feature information and the self-attention feature information.
[0014] In some embodiments, determining the retrieval results based on the cross-attention feature information and the self-attention feature information includes: determining the degree of matching between each candidate image and the image to be retrieved based on the dot product of the cross-attention feature information and the self-attention feature information; and determining the retrieval results based on the degree of matching.
[0015] In some embodiments, based on the feature vector of the image to be retrieved and the feature vector of each candidate image, the cross-attention model is used to determine the cross-attention feature information of the image to be retrieved, including: downsampling each candidate image to obtain a downsampled candidate image; based on the feature vector of the image to be retrieved and the feature vector of the downsampled candidate image, the cross-attention model is used to determine the cross-attention feature information.
[0016] In some embodiments, the first machine learning model consisting of the self-attention mechanism model and the cross-attention model is trained using a cross-entropy classification loss function.
[0017] In some embodiments, using the self-attention mechanism model, processing each candidate image in the candidate image set includes: downsampling each candidate image to obtain a downsampled candidate image; using the self-attention mechanism model, processing the feature vector of the downsampled candidate image to determine the self-attention feature information.
[0018] In some embodiments, determining a set of candidate images from a plurality of comparison images based on a degree of similarity between the image to be retrieved and each of the plurality of comparison images includes: extracting a feature vector of the image to be retrieved and a feature vector of each comparison image using a second machine learning model; and determining a set of candidate images based on a degree of similarity between the feature vector of the image to be retrieved and the feature vector of each comparison image.
[0019] In some embodiments, the second machine learning model is trained using a weighted mean of a center loss function and a cross entropy classification loss function.
[0020] According to other embodiments of the present disclosure, an image retrieval method is provided, comprising: determining a set of candidate images from a plurality of comparison images based on a degree of similarity between an image to be retrieved and each of a plurality of comparison images; and determining a retrieval result of the image to be retrieved from the set of candidate images based on a degree of similarity between the plurality of comparison images.
[0021] According to some further embodiments of the present disclosure, an image retrieval device is provided, comprising: a determination unit for determining a set of candidate images from a plurality of comparison images based on a degree of similarity between an image to be retrieved and each of the plurality of comparison images; and a retrieval unit for determining a retrieval result of an image to be retrieved from the set of candidate images based on a degree of similarity between the plurality of comparison images.
[0022] According to still further embodiments of the present disclosure, an image retrieval device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the image retrieval method in any of the above embodiments based on instructions stored in the memory device.
[0023] According to further embodiments of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the image retrieval method of any of the above embodiments is implemented.
[0024] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0026] FIG1 shows a flowchart of an image retrieval method according to some embodiments of the present disclosure;
[0027] FIG2 is a schematic diagram showing image transformation according to some embodiments of the present disclosure;
[0028] FIG3 shows a schematic diagram of preliminary image retrieval according to some embodiments of the present disclosure;
[0029] FIG4 shows a schematic diagram of a downsampling model according to some embodiments of the present disclosure;
[0030] FIG5 shows a schematic diagram of self-attention processing according to some embodiments of the present disclosure;
[0031] FIG6 shows a schematic diagram of a self-attention window according to some embodiments of the present disclosure;
[0032] FIG7 shows a schematic diagram of a cross-attention window according to some embodiments of the present disclosure;
[0033] FIG8 is a schematic diagram showing an image retrieval method according to some embodiments of the present disclosure;
[0034] FIG9 shows a flowchart of an image retrieval method according to some other embodiments of the present disclosure;
[0035] FIG10 shows a block diagram of an image retrieval apparatus according to some embodiments of the present disclosure;
[0036] FIG11 shows a block diagram of an image retrieval apparatus according to some other embodiments of the present disclosure;
[0037] FIG12 shows a block diagram of an image retrieval apparatus according to yet other embodiments of the present disclosure. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0039] Unless otherwise specifically stated, the relative arrangement of the parts and steps, the numerical expressions and the numerical values set forth in these embodiments do not limit the scope of the present disclosure. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to actual proportional relationships. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed herein, any specific values should be interpreted as being merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.
[0040] In image retrieval applications such as picture book image comparison, two images may differ by only a few lines of text. In this case, it's difficult to distinguish the two images based solely on features extracted by a CNN (Convolutional Neural Network). In other words, the features extracted by a CNN for similar images are highly similar, making it difficult to accurately identify the corresponding page in the picture book. Similarly, in person ID recognition, it's difficult to distinguish between two people wearing the same clothing.
[0041] In response to the above technical problems, the technical solution disclosed in the present invention can use the features extracted by CNN to perform preliminary image retrieval, obtain a set of all similar images to the image to be retrieved (such as the prob image), and then use the reRank technology to perform further screening in the image base library (such as a collection of gallery images).
[0042] On the other hand, during accurate image retrieval, a unique image that matches the image to be retrieved (which can be the same image as the image to be retrieved) is found in the image database. For example, image retrieval can be performed by detecting and matching feature points, that is, comparing the large number of details in two images to see if they are the same.
[0043] For example, in the process of accurate image retrieval, for a prob image, images with the same content as the prob image can be retrieved from the gallery. The retrieved images can be images of the same content as the prob image, but captured from different viewing angles, different lighting conditions, or different angles.
[0044] In some embodiments, a transformer model can be used to perform feature point matching to compare and match images. For example, among the multiple images initially retrieved, not only is the similarity between the prob image and the gallery image compared, but the similarity between the gallery images is also compared. This is similar to the image matching process performed by humans, which requires observing both the prob image and the similarity between the gallery images, thereby improving the accuracy of image retrieval.
[0045] In some embodiments, image retrieval can be achieved through two stages: preliminary retrieval and precise retrieval. First, CNN can be used to extract image features from the prob image. For example, CNN can include a ResNet (residual network) 34 model. Then, multiple gallery images that are most similar to the prob image (such as the top 10 in terms of similarity) are obtained. Subsequently, methods such as Rerank can be used to obtain multiple gallery images that are most similar to each of these gallery images in the gallery image set (such as the top 5 in terms of similarity); the union of the gallery images obtained twice can be used as the result of the preliminary retrieval.
[0046] In some embodiments, the transformer's self-attention mechanism and cross-attention mechanism can be used to determine the precise retrieval result by comprehensively considering the information between the features of each preliminary retrieval result and the information between the features of the prob image and the features of each preliminary retrieval result.
[0047] In this way, by expanding the gallery images that are similar to the prob image obtained for the first time, the technical problem of gallery images matching the prob image being incorrectly screened out due to the lack of accurate discrimination of features extracted by networks such as CNN can be solved, thereby improving the retrieval accuracy.
[0048] The inventors of the present disclosure have discovered that the above-mentioned related technologies have the following problem: the image features used as the retrieval basis are not accurate enough, resulting in a decrease in retrieval precision. In view of this, the present disclosure proposes an image retrieval technology solution that can improve retrieval precision.
[0049] For example, the following embodiments may be used to implement the technical solutions of the present disclosure.
[0050] FIG1 shows a flowchart of an image retrieval method according to some embodiments of the present disclosure.
[0051] As shown in FIG1 , in step S110, a candidate image set is determined from the plurality of comparison images based on the degree of similarity between the image to be retrieved and each of the plurality of comparison images. For example, the image to be retrieved may be a prob image, and the comparison images may be gallery images or images in an image base.
[0052] In some embodiments, the image to be retrieved may be perspective transformed to obtain multiple transformed images to be retrieved; and a candidate image set may be determined from the multiple compared images based on the degree of similarity between each transformed image to be retrieved and each compared image.
[0053] FIG2 shows a schematic diagram of image transformation according to some embodiments of the present disclosure.
[0054] As shown in Figure 2, the image to be retrieved can be transformed into transformed images from different perspectives, such as images a, b, and c. This ensures that images with the same content (such as the same object or person) from different perspectives can be effectively retrieved, thereby improving the accuracy of image retrieval.
[0055] In some embodiments, a second machine learning model is used to extract a feature vector of the image to be retrieved and a feature vector of each comparison image; and a set of candidate images is determined based on the degree of similarity between the feature vector of the image to be retrieved and the feature vector of each comparison image. For example, the second machine learning model is trained using a weighted mean of a center loss function and a cross entropy classification loss function.
[0056] In some embodiments, each comparison image can be treated as an image type, and a second machine learning model can be trained using a center loss function and a cross entropy classification loss function to classify the images to be retrieved. For example, the second machine learning model can include a CNN model.
[0057] For example, we can use the center loss function L C and the cross entropy classification loss function L S The weighted mean of determines the comprehensive loss function L:
[0058] i and j are the labels of image types, M is the number of image types, y i is the i-th image type, x i is the feature vector of the image to be matched, For image type y i The central eigenvector of , λ (e.g., the value can be between 0 and 1) is the weight used to adjust the importance of the two loss functions, and is a learnable parameter. For example, λ is used to balance the importance of the two loss functions and can be set to 1.0 to minimize the distance between images of the same category.
[0059] For example, when paying more attention to the differences between images of the same image type, the weight of the center loss function can be increased by adjusting λ; when paying more attention to the differences between images of different image types, the weight of the cross entropy classification loss function can be increased by adjusting λ.
[0060] In this way, the center loss function can better reflect the differences between images of the same image type, so as to improve the resolution ability of the machine learning model; the cross entropy classification loss function can better reflect the differences between images of different image types, so as to reduce the overfitting problem of the machine learning model; combining the two can complement each other and thus improve the performance of the machine learning model.
[0061] In some embodiments, ResNet50 can be used as a feature extraction network to obtain image features; then, a loss function L is used for reverse update. In the above embodiment, considering that the image differences within the same image type may be small, CenterLoss (center loss function) is used for training to improve the training effect; in addition, combined with ID loss, softmax and cross entropy classification loss functions can be used for training to improve the training effect.
[0062] In the above embodiment, a loss function combining softmax, cross entropy classification loss function, and center loss function is used. In this way, based on the preliminary distinction of image types by softmax, combined with the assistance of center loss, it is ensured that image features are not only distinguishable but also discriminable.
[0063] For example, based on the degree of similarity between each candidate image and the comparison image, at least one similar image corresponding to each candidate image can be determined from multiple comparison images; based on each candidate image and its at least one similar image, the candidate image set can be expanded; and the retrieval results can be determined from the expanded candidate image set.
[0064] In this way, through the expansion process, it can be ensured that the correct search results are included in the preliminary search results to improve the search accuracy.
[0065] For example, the above technical solution can be implemented through the embodiment in FIG3 .
[0066] FIG3 shows a schematic diagram of preliminary image retrieval according to some embodiments of the present disclosure.
[0067] As shown in Figure 3, multiple (e.g., 10) gallery images P1-P4 and N1-N6 with the highest similarity to the prob image (the shaded rectangle in Figure 3) are retrieved from all gallery images and serve as candidate image set A. Multiple (e.g., 5) gallery images that are most similar to each gallery image in candidate image set A are retrieved from all gallery images and serve as image sets B corresponding to each gallery image in candidate image set A.
[0068] In some embodiments, the number of first images belonging to the candidate image set among at least one similar image corresponding to any candidate image is calculated; and based on the first number of images, it is determined whether to expand the at least one similar image corresponding to any candidate image to the candidate image set.
[0069] For example, when the proportion of the number of first images in the number of second images is greater than a threshold (such as 60%, etc.), at least one similar image corresponding to any candidate image is expanded to the candidate image set, and the number of second images is the number of at least one similar image corresponding to any candidate image.
[0070] For example, if more than 60% of the gallery images in a certain image set B belong to the candidate image set A, the image set B (such as the image set B corresponding to images P1 to P4 in FIG3 ) is expanded to the candidate image set A. The expanded candidate image set A is used as the preliminary retrieval result.
[0071] In this way, it is fully guaranteed that the correct search results are included in the preliminary search results to improve the search accuracy. For example, based on the preliminary search results, the remaining steps in Figure 1 can be used to perform precise search to retrieve the correct search results.
[0072] As shown in FIG. 1 , in step S120 , a retrieval result of an image to be retrieved is determined from a set of candidate images based on the similarity between the multiple compared images.
[0073] In some embodiments, the similarity between the multiple compared images includes the similarity between each candidate image in the candidate image set and each compared image, and / or the similarity between multiple candidate images in the candidate image set.
[0074] In some embodiments, each candidate image is downsampled to obtain a downsampled candidate image; the feature vector of the downsampled candidate image is processed using a self-attention mechanism model to determine self-attention feature information; and the cross-attention feature information is determined using a cross-attention model based on the feature vector of the image to be retrieved and the feature vector of the downsampled candidate image.
[0075] In some embodiments, the degree of matching between each candidate image and the image to be retrieved is determined based on the dot product of the cross-attention feature information and the self-attention feature information; and the retrieval result is determined based on the degree of matching.
[0076] In this way, by extracting self-attention feature information and cross-attention feature information, the relevant information between images in the preliminary retrieval results and the relevant information between the image to be retrieved and the images in the preliminary retrieval results are fully considered, thereby improving the accuracy of the retrieval.
[0077] In some embodiments, a residual convolutional neural network can be used to downsample the preliminary search results. Images in the preliminary search results can be downsampled to 1 / 8 their original size to improve image retrieval efficiency and accuracy. For example, downsampling can be performed using the embodiment shown in FIG4 .
[0078] FIG4 shows a schematic diagram of a downsampling model according to some embodiments of the present disclosure.
[0079] As shown in Figure 4, the residual convolutional neural network used for downsampling processing may include a Relu (Linear rectification function) layer, a BatchNorm (batch normalization) layer, a 3×3 Con (convolution) layer with a stride of 2, a concat (connection) layer, and a 3×3 Con layer with a stride of 1. Through the above residual convolutional neural network, the images in the preliminary retrieval results are downsampled multiple times (e.g., 4 times), and each processing can downsample the image by a factor of 2.
[0080] For example, for an image with a resolution of 640×480, through four downsampling processes of the above residual convolutional neural network, a feature image of 80×60×64 can be obtained, where 64 is the number of image channels. Then, the feature image can be flattened to convert each image into a feature vector of 80×60×64. After adding the position code of the corresponding image to the feature vector, it is mapped to a 512-dimensional feature vector through a fully connected network, which serves as the input token for self-attention and cross-attention.
[0081] In some embodiments, a self-attention mechanism model is used to process each candidate image in the candidate image set to determine the self-attention feature information of each candidate image; and the retrieval result is determined based on the self-attention feature information of each candidate image and the feature information of the image to be retrieved.
[0082] In this way, through the self-attention feature information, the relevant information between the images in the preliminary retrieval results can be extracted as the basis for accurate retrieval, thereby improving the accuracy of the retrieval.
[0083] In some embodiments, a self-attention mechanism model is used to extract self-attention feature information based on feature vectors of multiple candidate images located in a self-attention window in a feature vector sequence, where the feature vector sequence includes a feature vector for each candidate image; the self-attention window is slid across the feature vector sequence based on a sliding step size; and the above extraction and sliding steps are repeated until feature vectors for all candidate images in the feature vector sequence are processed. For example, the self-attention mechanism model can be a multi-head self-attention mechanism model.
[0084] In this way, the self-attention window can be used to extract only a small number of self-attention information between feature vectors at a time, thereby improving the efficiency and accuracy of image retrieval.
[0085] In some embodiments, the length of the self-attention window is greater than the sliding step size.
[0086] In this way, by sliding the self-attention window, the objects processed each time have overlapping parts, which can maintain the relevance between each processing. Therefore, on the basis of ensuring that fewer feature vectors are processed each time, it can also ensure that
[0087] For example, the self-attention feature information may be processed through the embodiment in FIG5 .
[0088] FIG5 shows a schematic diagram of self-attention processing according to some embodiments of the present disclosure.
[0089] As shown in Figure 5, the length of the self-attention window marked by the dotted box is 4, and the sliding step size is 2. First, the self-attention feature information within the current self-attention window is calculated; then, the self-attention window is slid, and the self-attention feature information calculation is continued at the position after the sliding.
[0090] In this way, it is possible to consider the information between feature vectors during the precise retrieval process, thereby improving the accuracy of image retrieval.
[0091] FIG6 shows a schematic diagram of a self-attention window according to some embodiments of the present disclosure.
[0092] [Corrected 08.11.2023 according to Rule 91] As shown in Figure 6, the self-attention mechanism module includes MLP (Multilayer Perceptron) layer, Layer Norm (normalization layer), Add processing module, Windows Multi Head Attention (multi-head attention mechanism window), Shift Windows Multi Head Attention (sliding multi-head attention mechanism window), etc.
[0093] In some embodiments, the cross-attention feature information of the image to be retrieved is determined based on the feature vector of the image to be retrieved and the feature vector of each candidate image using a cross-attention model; and the retrieval result is determined based on the cross-attention feature information and the self-attention feature information.
[0094] In this way, not only the information between the candidate images (ie, the difference information) is considered, but also the relevant information between the retrieved image and the candidate images is considered, thereby improving the accuracy of image retrieval.
[0095] For example, cross-attention feature information can be extracted through the embodiment in Figure 7.
[0096] FIG7 shows a schematic diagram of a cross-attention window according to some embodiments of the present disclosure.
[0097] As shown in Figure 7, the feature vectors of the candidate images are weighted through the self-attention mechanism (i.e., the update process in Figure 7) to obtain self-attention feature information, i.e., the q (query), k (key), and v (value) feature vectors of the candidate images. On this basis, the cross-attention feature information is calculated based on the q, k, and v feature vectors of the image to be retrieved and the q, k, and v feature vectors of each candidate image.
[0098] For example, the self-attention mechanism model and the cross-attention mechanism model can be used as a processing module and repeated multiple times (such as 4 times) to obtain the feature vector of the transformed image to be retrieved and the feature vector of the image in the preliminary retrieval result; then, the dot product of the feature vector of the image to be retrieved and the feature vector of the image in the preliminary retrieval result is calculated; the dot product is processed by softmax to calculate the similarity between the image in each preliminary retrieval result and the image to be retrieved.
[0099] In this way, not only the difference information between the candidate images is considered, but also the related information between the retrieved image and the candidate images is considered, thereby improving the accuracy of image retrieval.
[0100] In some embodiments, the first machine learning model consisting of the self-attention mechanism model and the cross-attention model is trained using a cross-entropy classification loss function.
[0101] For example, the cross entropy classification loss function L CS It can be:
[0102] are learnable parameters, is the feature vector.
[0103] FIG8 shows a schematic diagram of an image retrieval method according to some embodiments of the present disclosure.
[0104] As shown in Figure 8, in the preliminary search, CNN is first used to extract features from the prob image and gallery images. For example, ResNet34 can be used for feature extraction. The cosine distance between the feature vector of the prob image and the feature vector of the gallery image is calculated to identify the 10 closest gallery images. Then, using Rerank processing, the distance between the eigenvectors and the 5 gallery images closest to each of the 10 gallery images is calculated. All retrieved images are then combined as the preliminary search results.
[0105] In precise retrieval, CNN is used again to extract local features from the results of the preliminary retrieval to generate the corresponding feature map, with its length and width reduced to 1 / 8 of the original and the channel direction to 64; then, the feature map is expanded and linearly mapped to generate a 2048 feature vector.
[0106] For example, a self-attention mechanism can be implemented based on windows and sliding windows. First, the attention mechanism calculation within the window is implemented (i.e., W-MSA). For example, a multi-head attention mechanism based on the window can be used for calculation. Then, the attention mechanism calculation based on the sliding window (i.e., SW-MSA) is performed to ensure information exchange between windows, thereby realizing information exchange between various features.
[0107] For example, a cross-attention mechanism can be calculated based on the feature vectors of the prob image and the gallery image. The self-attention mechanism calculation above captures information about the initial search results in the underlying database. Therefore, a cross-attention mechanism can be calculated for the feature vectors of the prob image and the feature vectors of the initial search results.
[0108] For example, the self-attention mechanism calculation and the cross-attention mechanism calculation can be combined to form a block calculation module, and the calculation can be repeated three times.
[0109] For example, by calculating the dot product of the feature vector of the prob image and the feature vector of the preliminary search result, combined with softmax processing, the similarity between each preliminary search result and the prob image can be obtained. The one with the highest similarity can be used as the search result.
[0110] In some embodiments, for pedestrian identification applications, identity information can be pre-associated with human images (i.e., comparison images) in the image database. After obtaining a pedestrian image (i.e., the image to be retrieved), the disclosed technical solution can be used to accurately find a comparison image that matches the pedestrian image, and the identity information associated with the comparison image can be obtained. This allows for rapid and accurate pedestrian identification.
[0111] For example, a pedestrian image can be transformed into images to be retrieved from different perspectives. This ensures that the same pedestrian can be effectively identified from images from different perspectives, thereby improving the accuracy of image retrieval.
[0112] In some embodiments, for the application scenario of picture book image matching, an explanation media file can be matched in advance for the comparison image in the image base library; after the user selects the current page in the picture book, the technical solution of the present disclosure is used to accurately find the comparison image matching the current page; and the explanation media file corresponding to the comparison image is played, so that the user can quickly and accurately understand the relevant content of the current page, thereby improving the user experience.
[0113] FIG9 shows a flowchart of an image retrieval method according to some other embodiments of the present disclosure.
[0114] As shown in Figure 9, in the picture book image matching scenario, after turning to a page of the current picture book (i.e., the image to be searched), it is necessary to quickly find which page of which picture book the current page belongs to (i.e., the search result) in the image base library (i.e., multiple comparison images). In step S210, a preliminary search is performed to obtain the preliminary search results for the current page, i.e., the candidate image set.
[0115] In step S220, the proportion of multiple candidate images in the candidate image set in each picture book in the base library is determined to determine which picture book the current page belongs to; ReRank processing is performed to expand the results of the initial search.
[0116] In image retrieval applications such as picture book image comparison, two images may differ by only a few lines of text. In this case, it's difficult to distinguish the two images based solely on features extracted by a CNN. Specifically, the features extracted by a CNN for similar images are highly similar, making it difficult to accurately identify the page in the picture book that corresponds to the current page. In this case, the candidate image set determined solely based on CNN features is likely to not include an image that matches the current page.
[0117] To address these technical issues, we use ReRank technology to further filter the image base based on the initially selected candidate images to expand the initial search results. This prevents images that match the current page from being mistakenly filtered out during the initial screening process, thereby improving the accuracy of image matching.
[0118] In step 230, a pre-verification removal process is performed. For example, images that do not belong to the current picture book can be removed from the results of the preliminary search of the current page to serve as input data for the precise search.
[0119] In step S240, the feature vector of the result of the preliminary retrieval after the removal process is extracted and calculated using the self-attention mechanism and the cross-attention mechanism.
[0120] In this way, the information between the preliminary search results and the information between the prob image and the preliminary search results are fully considered, thereby improving the retrieval accuracy.
[0121] In step S250, the retrieval result is determined by comparing the similarity (such as the score map of each candidate image).
[0122] In step S260, the search results are output.
[0123] FIG10 shows a block diagram of an image retrieval apparatus according to some embodiments of the present disclosure.
[0124] As shown in FIG10 , the image retrieval device 10 includes: a determination unit 101 for determining a candidate image set from a plurality of comparison images based on a degree of similarity between the image to be retrieved and each of the plurality of comparison images; and a retrieval unit 102 for determining a retrieval result of the image to be retrieved from the candidate image set based on a degree of similarity between the plurality of comparison images.
[0125] In some embodiments, the similarity between the multiple compared images includes the similarity between each candidate image in the candidate image set and each compared image, and / or the similarity between multiple candidate images in the candidate image set.
[0126] In some embodiments, the retrieval unit 102 determines at least one similar image corresponding to each candidate image from multiple comparison images based on the degree of similarity between each candidate image and the comparison image; expands the candidate image set based on each candidate image and its at least one similar image, and determines the retrieval result from the expanded candidate image set.
[0127] In some embodiments, the retrieval unit 102 calculates the number of first images belonging to the candidate image set among at least one similar image corresponding to any candidate image, and determines whether to expand the at least one similar image corresponding to any candidate image to the candidate image set based on the first number of images.
[0128] In some embodiments, when the proportion of the first image number in the second image number is greater than a threshold, the retrieval unit 102 expands at least one similar image corresponding to any candidate image to the candidate image set, and the second image number is the number of at least one similar image corresponding to any candidate image.
[0129] In some embodiments, the retrieval unit 102 uses a self-attention mechanism model to process each candidate image in the candidate image set to determine the self-attention feature information of each candidate image; and determines the retrieval result based on the self-attention feature information of each candidate image and the feature information of the image to be retrieved.
[0130] In some embodiments, the retrieval unit 102 extracts self-attention feature information using a self-attention mechanism model based on the feature vectors of multiple candidate images located in the self-attention window in the feature vector sequence, where the feature vector sequence includes the feature vector of each candidate image; slides the self-attention window on the feature vector sequence according to the sliding step size; and repeats the above-mentioned extraction step and sliding step until the feature vectors of all candidate images in the feature vector sequence are processed.
[0131] In some embodiments, the length of the self-attention window is greater than the sliding step size.
[0132] In some embodiments, the retrieval unit 102 uses a cross-attention model to determine the cross-attention feature information of the image to be retrieved based on the feature vector of the image to be retrieved and the feature vector of each candidate image, and determines the retrieval result based on the cross-attention feature information and the self-attention feature information.
[0133] In some embodiments, the retrieval unit 102 determines the degree of matching between each candidate image and the image to be retrieved based on the dot product of the cross-attention feature information and the self-attention feature information, and determines the retrieval result based on the degree of matching.
[0134] In some embodiments, the retrieval unit 102 downsamples each candidate image to obtain a downsampled candidate image, and determines the cross-attention feature information using a cross-attention model based on the feature vector of the image to be retrieved and the feature vector of the downsampled candidate image.
[0135] In some embodiments, the first machine learning model consisting of the self-attention mechanism model and the cross-attention model is trained using a cross-entropy classification loss function.
[0136] In some embodiments, the retrieval unit 102 downsamples each candidate image to obtain a downsampled candidate image, and uses a self-attention mechanism model to process the feature vector of the downsampled candidate image to determine self-attention feature information.
[0137] In some embodiments, the determination unit 101 uses a second machine learning model to extract a feature vector of the image to be retrieved and a feature vector of each comparison image; and determines a set of candidate images based on the degree of similarity between the feature vector of the image to be retrieved and the feature vector of each comparison image.
[0138] In some embodiments, the second machine learning model is trained using a weighted mean of a center loss function and a cross entropy classification loss function.
[0139] FIG11 shows a block diagram of an image retrieval apparatus according to some other embodiments of the present disclosure.
[0140] As shown in FIG11 , the image retrieval system 11 of this embodiment includes a memory 111 and a processor 112 coupled to the memory 111 . The processor 112 is configured to execute the image retrieval method of any one of the embodiments of the present disclosure based on instructions stored in the memory 111 .
[0141] The memory 111 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, a database, and other programs.
[0142] FIG12 shows a block diagram of an image retrieval apparatus according to yet other embodiments of the present disclosure.
[0143] As shown in FIG12 , the image retrieval device 12 of this embodiment includes: a memory 1210 and a processor 1220 coupled to the memory 1210 . The processor 1220 is configured to execute the method in any one of the aforementioned embodiments based on instructions stored in the memory 1210 .
[0144] The memory 1210 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, application programs, a boot loader, and other programs.
[0145] The image retrieval device 12 may also include an input / output interface 1230, a network interface 1240, a storage interface 1250, and the like. These interfaces 1230, 1240, and 1250, as well as the memory 1210 and the processor 1220, may be connected, for example, via a bus 1260. The input / output interface 1230 provides a connection interface for input / output devices such as a display, mouse, keyboard, touch screen, microphone, and speakers. The network interface 1240 provides a connection interface for various networked devices. The storage interface 1250 provides a connection interface for external storage devices such as SD cards and USB flash drives.
[0146] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0147] So far, the present invention has been described in detail. In order to avoid obscuring the concept of the present invention, some details well known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed herein.
[0148] The methods and systems of the present disclosure may be implemented in many ways. For example, the methods and systems of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0149] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. An image retrieval device, comprising at least one processor, wherein the at least one processor is configured to: Determine a candidate image set from the multiple compared images according to the similarity between the image to be retrieved and each of the multiple compared images; According to the similarity between the multiple compared images, a retrieval result of the image to be retrieved is determined from the candidate image set.
2. The image retrieval device according to claim 1, wherein: The similarity between the multiple compared images includes the similarity between each candidate image in the candidate image set and the compared image, and / or the similarity between multiple candidate images in the candidate image set.
3. The image retrieval device according to claim 2, wherein: Determining the retrieval result of the image to be retrieved from the candidate image set according to the similarity between the multiple compared images includes: Determining at least one similar image corresponding to each candidate image from the plurality of compared images according to the degree of similarity between each candidate image and the compared image; Expanding the candidate image set according to each candidate image and at least one similar image thereof; The retrieval result is determined from the expanded candidate image set.
4. The image retrieval device according to claim 3, wherein: The expanding the candidate image set according to each candidate image and the at least one similar image comprises: Calculating the number of first images belonging to the candidate image set in at least one similar image corresponding to any candidate image; According to the first number of images, it is determined whether to expand at least one similar image corresponding to any one of the candidate images to the candidate image set.
5. The image retrieval device according to claim 4, wherein: The determining, according to the number of images, whether to expand at least one similar image corresponding to any one of the candidate images to the candidate image set comprises: When the proportion of the first number of images in the second number of images is greater than a threshold, at least one similar image corresponding to any one of the candidate images is expanded to the candidate image set, and the second number of images is the number of at least one similar image corresponding to any one of the candidate images.
6. The image retrieval device according to any one of claims 1 to 5, wherein: Determining the retrieval result of the image to be retrieved from the candidate image set according to the similarity between the multiple compared images includes: Using a self-attention mechanism model, processing each candidate image in the candidate image set to determine self-attention feature information of each candidate image; The retrieval result is determined according to the self-attention feature information of each candidate image and the feature information of the image to be retrieved.
7. The image retrieval device according to claim 6, wherein: The using the self-attention mechanism model to process each candidate image in the candidate image set to determine the self-attention feature information of each candidate image includes: Extracting the self-attention feature information using the self-attention mechanism model according to the feature vectors of multiple candidate images located in the self-attention window in the feature vector sequence, wherein the feature vector sequence includes the feature vector of each candidate image; Sliding the self-attention window on the feature vector sequence according to a sliding step size; The above-mentioned extraction step and sliding step are repeated until the feature vectors of all candidate images in the feature vector sequence are processed.
8. The image retrieval device according to claim 7, wherein: The length of the self-attention window is greater than the sliding step size.
9. The image retrieval device according to claim 6, wherein: The determining the retrieval result according to the self-attention feature information of each candidate image and the feature information of the image to be retrieved comprises: Determine cross-attention feature information of the image to be retrieved using a cross-attention model according to the feature vector of the image to be retrieved and the feature vector of each candidate image; The retrieval result is determined based on the cross-attention feature information and the self-attention feature information.
10. The image retrieval device according to claim 9, wherein: The determining the retrieval result according to the cross-attention feature information and the self-attention feature information comprises: Determining a matching degree between each candidate image and the image to be retrieved according to a dot product of the cross-attention feature information and the self-attention feature information; The search result is determined according to the matching degree.
11. The image retrieval device according to claim 9, wherein: The determining the cross-attention feature information of the image to be retrieved by using the cross-attention model according to the feature vector of the image to be retrieved and the feature vector of each candidate image includes: Performing downsampling processing on each of the candidate images to obtain a downsampled candidate image; The cross-attention feature information is determined using the cross-attention model according to the feature vector of the image to be retrieved and the feature vector of the down-sampled candidate image.
12. The image retrieval device according to claim 9, wherein: The first machine learning model consisting of the self-attention mechanism model and the cross-attention model is trained by a cross-entropy classification loss function.
13. The image retrieval device according to claim 6, wherein: The process of processing each candidate image in the candidate image set by using the self-attention mechanism model includes: Performing downsampling processing on each of the candidate images to obtain a downsampled candidate image; The self-attention mechanism model is used to process the feature vector of the downsampled candidate image to determine the self-attention feature information.
14. The image retrieval device according to any one of claims 1 to 5, wherein: The step of determining a candidate image set from the multiple compared images according to the similarity between the image to be retrieved and each of the multiple compared images comprises: Extracting a feature vector of the image to be retrieved and a feature vector of each of the compared images using a second machine learning model; The candidate image set is determined according to the similarity between the feature vector of the image to be retrieved and the feature vector of each of the comparison images.
15. The image retrieval device according to claim 14, wherein: The second machine learning model is trained using a weighted mean of a center loss function and a cross entropy classification loss function.
16. An image retrieval method, comprising: Determine a candidate image set from the multiple compared images according to the similarity between the image to be retrieved and each of the multiple compared images; According to the similarity between the multiple compared images, a retrieval result of the image to be retrieved is determined from the candidate image set.
17. An image retrieval device, comprising: a determination unit, configured to determine a candidate image set from among the plurality of comparison images according to a degree of similarity between the image to be retrieved and each of the plurality of comparison images; The retrieval unit is used to determine the retrieval result of the image to be retrieved from the candidate image set according to the similarity between the multiple comparison images.
18. An image retrieval device, comprising: Memory; and A processor coupled to the memory, wherein the processor is configured to execute the image retrieval method of claim 16 based on instructions stored in the memory device.
19. A non-volatile computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the image retrieval method according to claim 16 is implemented.