Image Retrieval Method Based on Bag-of-Words Model Constructed by Hybrid KAZE Algorithm

The mixed KAZE algorithm extracts region and edge features and builds a hybrid visual word package, which solves the problem of insufficient accuracy of traditional image retrieval methods and achieves a more accurate image retrieval effect.

CN115620037BActive Publication Date: 2025-06-20CHINESE PEOPLES LIBERATION ARMY UNIT 63816
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211385212.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-06-20
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

The traditional image retrieval method based on word packet model has insufficient image retrieval accuracy, and cannot effectively describe the complete features of the image, resulting in inaccurate similarity.

Method used

The hybrid KAZE algorithm is used to extract regional features and edge features, build regional visual word packages and edge visual word packages, and generate mixed visual word packages through parallel operations for subsequent feature matching and post-processing.

Benefits of technology

By combining region and edge features, more complete image features are generated, the accuracy and accuracy of image retrieval is improved and the reflection of image similarity is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620037B_ABST
    Figure CN115620037B_ABST
Patent Text Reader

Abstract

This application relates to an image retrieval method for constructing a bag-of-words model based on the hybrid KAZE algorithm. By using a pre-set feature extraction algorithm to extract the feature vectors of regions and edges in the image respectively, then performing feature quantization and constructing indexes on the feature vectors of the two respectively to obtain the visual bag-of-words of the two. At the same time, vectorizing the visual bag-of-words of the two respectively to obtain the bag-of-words vectors of the two, and juxtaposing the visual bag-of-words of the two to obtain a hybrid visual bag-of-words. At the same time, also performing feature matching on the bag-of-words vectors of the two respectively to obtain the similarity of the two, calculating the hybrid similarity according to the similarity of the two, then sorting the images in the image database in descending order according to the hybrid similarity to obtain the initial image retrieval result, and finally performing post-processing on the initial image retrieval result in combination with the hybrid visual bag-of-words to obtain the final image retrieval result. Using this method can improve the overall accuracy of image retrieval and effectively improve the image retrieval accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to an image retrieval method based on a bag-of-words model constructed by a hybrid KAZE algorithm. Background Art

[0002] With the rise of the Internet, the picture data on the network has grown at an astonishing speed, forming a huge image retrieval database. For the massive picture data, how to obtain the pictures required by users through image retrieval has become a technical problem to be solved at present.

[0003] In traditional technologies, the most common image retrieval draws on the idea of text retrieval, analogizes an image to a piece of text, the local features extracted from the image correspond to the words in the text, uses the frequency of the words appearing to establish a bag-of-words vector to represent the image, and retrieves the image by calculating and comparing the similarity between the bag-of-words vectors. This method is called the image retrieval method based on the bag-of-words model, and mainly includes the following five steps: feature extraction, feature quantization, index construction, feature matching, and post-processing.

[0004] However, in the process of implementing the present invention, the inventor found that the aforementioned traditional image retrieval method still has the technical problem of low image retrieval accuracy. Summary of the Invention

[0005] Based on the problem of low accuracy of the above-mentioned traditional image retrieval technology, an image retrieval method based on a bag-of-words model constructed by a hybrid KAZE algorithm is provided to solve the above problem.

[0006] An image retrieval method based on a bag-of-words model constructed by a hybrid KAZE algorithm includes:

[0007] Extracting the region feature vector of the reference image by using a pre-set region feature extraction algorithm, and extracting the edge feature vector of the reference image by using a pre-set edge feature extraction algorithm; all the region feature vectors of all the reference images in the image database are used to generate a region visual dictionary, and all the edge feature vectors are used to generate an edge visual dictionary;

[0008] Respectively performing feature quantization and index construction on the region feature vector and the edge feature vector to obtain a region visual bag-of-words and an edge visual bag-of-words, and vectorizing the region visual bag-of-words and the edge visual bag-of-words to obtain a region visual bag-of-words vector and an edge visual bag-of-words vector; the visual bag-of-words includes a visual word index, a feature position, and a visual dictionary capacity, and the visual bag-of-words is a region visual bag-of-words or an edge visual bag-of-words;

[0009] Performing a juxtaposition operation on the region visual bag-of-words and the edge visual bag-of-words to obtain a hybrid visual bag-of-words and storing the hybrid visual bag-of-words in a post-processing program;

[0010] Perform feature matching on the regional visual word bag vector and the edge visual word bag vector respectively to obtain the regional similarity and the edge similarity;

[0011] Take the square root of the sum of the square of the edge similarity and the square of the regional similarity to obtain the hybrid similarity;

[0012] Arrange all the reference images in the image database in descending order according to the hybrid similarity to obtain the initial image retrieval result;

[0013] Perform post-processing based on the initial image retrieval result in combination with the hybrid visual word bag to obtain the final image retrieval result.

[0014] In one embodiment, the process of juxtaposing operations includes:

[0015] Perform a juxtaposing operation on the visual word indexes in the regional visual word bag and the visual word indexes in the edge visual word bag to form the hybrid visual word index of the hybrid visual word bag;

[0016] Perform a juxtaposing operation on the feature positions in the regional visual word bag and the feature positions in the edge visual word bag to form the hybrid feature position of the hybrid visual word bag;

[0017] Perform a summation operation on the visual dictionary capacities in the regional visual word bag and the edge visual word bag to form the hybrid visual dictionary capacity;

[0018] The hybrid visual word index, the hybrid feature position and the hybrid visual dictionary capacity form the hybrid visual word bag.

[0019] In one embodiment, the hybrid visual word bag includes (n + m) hybrid visual word indexes, (n + m) hybrid feature positions and the hybrid visual dictionary capacity, where n is the number of regional features, m is the number of edge features, and the size of the hybrid visual dictionary capacity is 2Num.

[0020] In one embodiment, the regional visual word bag includes n regional visual word indexes, n regional feature positions and the regional visual dictionary capacity, where the size of the regional visual dictionary capacity is Num.

[0021] In one embodiment, the edge visual word bag includes m edge visual word indexes, m edge feature positions and the edge visual dictionary capacity, where the size of the edge visual dictionary capacity is Num.

[0022] In one embodiment, the process of feature quantization and index construction includes:

[0023] Cluster the regional feature vectors into regional visual words according to the vector similarity;

[0024] Cluster the edge feature vectors into edge visual words according to the vector similarity;

[0025] Merge the regional visual words and the edge visual words to generate a visual dictionary;

[0026] Establish a one-to-one correspondence between the image features of the image and the visual words in the visual dictionary to generate a visual word index, and establish a one-to-one correspondence between the visual words in the visual dictionary and the image features of the image to generate an inverted index;

[0027] Locate the position of the image features in the image to obtain the feature positions, and use the visual word index, the feature positions, and the visual dictionary capacity to form a visual word bag.

[0028] In one embodiment, the vectorization process includes:

[0029] Extract the visual word index of the visual word bag;

[0030] Arrange according to the frequency and order of appearance of the visual words in the visual word index;

[0031] Obtain a visual word bag vector; the visual word bag vector includes a regional visual word bag vector and an edge visual word bag vector.

[0032] In one embodiment, the feature matching process includes:

[0033] Calculate the cosine distance between the regional visual word bag vector of the query image and the regional visual word bag vectors of all reference images in the image database to obtain a regional cosine distance, and use 1 minus the regional cosine distance to obtain the regional similarity between the query image and all reference images;

[0034] Calculate the cosine distance between the edge visual word bag vector of the query image and the edge visual word bag vectors of all reference images in the image database to obtain an edge cosine distance, and use 1 minus the edge cosine distance to obtain the edge similarity between the query image and all reference images.

[0035] In one embodiment, the post-processing process includes:

[0036] Extract the top K reference images with the highest to lowest mixed similarity in the initial image retrieval results to obtain a first set of reference images;

[0037] Use the feature positions and visual word index in the mixed visual word bag to perform a secondary comparison between the reference images in the first set of reference images and the query image;

[0038] Generate a new similarity metric through the secondary comparison, and perform a descending order again according to the new similarity metric to obtain the final image retrieval result.

[0039] It should be noted that the pre-set region feature extraction algorithm and the pre-set edge feature extraction algorithm are the hybrid KAZE algorithm.

[0040] The above image retrieval method for constructing a bag-of-words model based on the hybrid KAZE algorithm extracts the region feature vectors in the image by using the pre-set region feature extraction algorithm, extracts the edge feature vectors in the image by using the pre-set edge feature extraction algorithm, and then performs feature construction and quantization indexing on the extracted image feature vectors to obtain the region visual word bag and the edge visual word bag. Then, the region visual word bag and the edge visual word bag are vectorized to obtain the region word bag vector and the edge word bag vector. The region visual word bag and the edge visual word bag are juxtaposed to obtain the hybrid visual word bag, which is used in the subsequent verification process of post-processing. At the same time, the region word bag vector and the edge word bag vector are also respectively subjected to feature matching to obtain the region similarity and the edge similarity. The hybrid similarity is calculated based on the region similarity and the edge similarity. Then, the reference images in the image database are sorted in descending order according to the hybrid similarity to obtain the initial image retrieval result. Finally, in the post-processing, the initial image retrieval result is further filtered and re-sorted in combination with the hybrid visual word bag to obtain the final image retrieval result. By juxtaposing the region visual word bag and the edge visual word bag to obtain the hybrid visual word bag, more complete and comprehensive image features are obtained, which are used in the subsequent verification process of post-processing, making the final image retrieval result after verification more accurate. At the same time, by calculating the region similarity and the edge similarity to obtain the hybrid similarity, a more complete and accurate initial image retrieval result is obtained, thereby overall improving the overall accuracy of image retrieval and effectively improving the image retrieval accuracy. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] Figure 1 It is a schematic flowchart of the method for an image retrieval method based on the hybrid KAZE algorithm for constructing a bag-of-words model in an embodiment;

[0043] Figure 2 It is a schematic module structure diagram of an image retrieval device based on the hybrid KAZE algorithm for constructing a bag-of-words model in an embodiment. Detailed Embodiments

[0044] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.

[0045] It should be noted that referring to "embodiment" in this document means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase is shown at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0046] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0047] In the research of the present application, the inventors found that when using the image retrieval method based on the bag-of-words model in traditional image retrieval technology, in the feature extraction stage, generally used feature detection and extraction algorithms such as SIFT, SURF, BRISK or ORB, etc. Among them, SIFT and SURF algorithms are for blob features, while BRISK and ORB algorithms are for corner features. The local features extracted by them are relatively single, resulting in the visual bag-of-words constructed by local features not being able to comprehensively describe the overall image, so the similarity cannot fully reflect the similarity between the query image and the reference image. In post-processing, due to the incompleteness of the visual bag-of-words, the elimination of incorrect results and the return of correct results are inaccurate, ultimately reducing the accuracy of image retrieval.

[0048] In one embodiment, as Figure 1 shown, there is provided an image retrieval method for constructing a bag-of-words model based on a hybrid KAZE algorithm, including:

[0049] Step 102, extracting the region feature vector of the reference image by using a pre-set region feature extraction algorithm, and extracting the edge feature vector of the reference image by using a pre-set edge feature extraction algorithm; all region feature vectors of all reference images in the image database are used to generate a region visual dictionary, and all edge feature vectors are used to generate an edge visual dictionary.

[0050] The regional feature vector in this step refers to the aggregation of a part of the image features in the part of the image area with a larger width in the retrieved image; the edge local feature refers to the aggregation of a part of the image features in the part of the image area near the boundary. This part of the features is often easily overlooked, resulting in inaccurate image retrieval. In this step, it is used as an equally important reference with the regional local feature, which can avoid the problem of inaccurate overall image retrieval due to incomplete image boundary information.

[0051] Step 104: Respectively perform feature quantization and construct an index on the regional feature vector and the edge feature vector to obtain a regional visual word bag and an edge visual word bag, and vectorize the regional visual word bag and the edge visual word bag to obtain a regional visual word bag vector and an edge visual word bag vector; the visual word bag includes a visual word index, a feature position, and a visual dictionary capacity, and the visual word bag is a regional visual word bag or an edge visual word bag.

[0052] In this step, the visual word index represents the visual word corresponding to this local feature, the feature position represents the spatial coordinates of this local feature on this image, the visual dictionary capacity represents the number of all visual words extracted from all images in an image database, and the visual word bag vector represents the image as a K-dimensional numerical vector based on counting the number of times each visual word appears in this image. By constructing the visual word bag and the visual word bag vector, making the features of the image correspond one by one with the visual word index can obtain the corresponding image retrieval result accurately and quickly when the image is retrieved, and arranging them in an ordered arrangement order can improve the retrieval efficiency.

[0053] Step 106: Perform a juxtaposition operation on the regional visual word bag and the edge visual word bag to obtain a mixed visual word bag, and store the mixed visual word bag in the post-processing program.

[0054] It should be noted that the juxtaposition operation in this step is to perform a juxtaposition operation on the visual word index, the feature position, and the visual dictionary capacity respectively, and then integrate the results after the juxtaposition operation to obtain a mixed visual word bag. This step expands the number of image features of the image, and due to the targeted feature extraction in step 102 and the more accurate retrieval process in step 104, the mixed visual word bag has a more comprehensive and accurate image feature capacity, so as to further filter out unmatched images in the subsequent post-processing process and re-rank to obtain the final retrieval result, so as to achieve the purpose of improving the accuracy of image retrieval.

[0055] Step 108: Respectively perform feature matching on the regional visual word bag vector and the edge visual word bag vector to obtain a regional similarity and an edge similarity.

[0056] In this step, let the query image be represented as Q, the reference image be represented as R, the cosine distance be represented as D, and the edge visual word bag vector be represented as Iedge , the regional visual word bag vector is represented as I region , then the cosine distance between the query image and the regional visual word bag vector of the reference image is represented as The cosine distance between the query image and the edge visual word bag vector of the reference image is represented as Then the regional similarity S region and the edge similarity S edge The calculation method is as follows:

[0057]

[0058]

[0059] In this step, by performing feature matching on the regional visual word bag vector and the edge visual word bag vector, the regional similarity and the edge similarity are obtained, providing accurate data for the calculation of the subsequent hybrid similarity.

[0060] Step 110, take the square root of the sum of the square of the edge similarity and the square of the regional similarity to obtain the hybrid similarity.

[0061] In this step, the edge similarity can be represented as S region , the hybrid similarity can be represented as S edge , the hybrid similarity can be represented as S mixed , and their relationship can be represented as:

[0062]

[0063] Step 112, sort all the reference images in the image database in descending order according to the hybrid similarity to obtain the initial image retrieval result.

[0064] In this step, the specific method of sorting in descending order is to compare all the reference images in the image database with the query image in order of higher hybrid similarity value first. After comparison, the reference images with a certain value are obtained, and then these reference images with a certain value are used as the initial image retrieval result and used as the input value in the subsequent post-processing stage. Sorting in descending order ensures the credibility of the initial image retrieval result.

[0065] Step 114, perform post-processing on the hybrid visual word bag based on the initial image retrieval result to obtain the final image retrieval result.

[0066] Verify the initial image retrieval result by combining the visual word index and feature position of the hybrid visual word bag, filter the unmatched images, and reorder to obtain the final image retrieval result. In this step, the initial image retrieval result is re-verified by the rich and accurate image features in the hybrid visual word bag, making the final image retrieval result more accurate.

[0067] In the above image retrieval method for constructing a bag-of-words model based on the hybrid KAZE algorithm, a region feature extraction algorithm set in advance is used to extract the region feature vector of the reference image, and an edge feature extraction algorithm set in advance is used to extract the edge feature vector of the reference image. Then, the obtained region feature vector and edge feature vector are subjected to feature quantization and index construction to obtain the corresponding visual bag of words and visual bag-of-words vector, and they are applied to the subsequent construction of the hybrid bag of words and the calculation of the hybrid similarity, so that the hybrid bag of words can more completely describe the image, and the hybrid similarity can more accurately reflect the similarity between the query image and the reference image. All the reference images are sorted in descending order of the hybrid similarity to obtain the initial image retrieval result. Based on the initial image retrieval result and combined with the hybrid visual bag of words, post-processing is performed to filter out unmatched images and re-rank to obtain the final image retrieval result, thereby improving the overall accuracy of image retrieval and effectively enhancing the image retrieval accuracy.

[0068] In one embodiment, the parallel operation process includes:

[0069] Performing a parallel operation on the visual word indexes in the region visual bag of words and the visual word indexes in the edge visual bag of words to form the hybrid visual word index of the hybrid visual bag of words;

[0070] Performing a parallel operation on the feature positions in the region visual bag of words and the feature positions in the edge visual bag of words to form the hybrid feature position of the hybrid visual bag of words;

[0071] Performing a summation operation on the visual dictionary capacities in the region visual bag of words and the edge visual bag of words to form the hybrid visual dictionary capacity;

[0072] The hybrid visual word index, hybrid feature position, and hybrid visual dictionary capacity constitute the hybrid visual bag of words.

[0073] Among them, the visual word index WordIndex, feature position Location, and visual dictionary capacity VocabularySize of the local feature can constitute a visual bag of words of an image feature. It should be noted that the region visual bag of words, edge visual bag of words, and hybrid visual bag of words all belong to the visual bag of words containing local features. In this embodiment, the hybrid visual bag of words is obtained by performing a parallel operation on the region visual bag of words and the edge visual bag of words, which not only ensures the arrangement order of the visual word index and feature position, but also expands the number of the visual word index and feature position itself, so as to more accurately filter out unmatched image retrieval results in the subsequent post-processing process.

[0074] In one embodiment, the hybrid visual word bag includes (n + m) hybrid visual word indices, (n + m) hybrid feature positions, and a hybrid visual dictionary capacity, where n is the number of regional features and m is the number of edge features, and the size of the hybrid visual dictionary capacity is 2Num.

[0075] Among them, the hybrid visual word index can be expressed as The hybrid feature position index can be expressed as The hybrid visual dictionary capacity can be expressed as 2Num. Among them, Num in the hybrid visual word index is intended to represent that after the juxtaposition operation, the previous visual words are corresponding in the hybrid visual dictionary to ensure the correctness of the program during execution. It should be noted that in this embodiment, the arrangement order and quantity of the hybrid visual word index and the hybrid feature position are defined to ensure the accuracy of the hybrid visual word bag.

[0076] In one embodiment, the regional visual word bag includes n regional visual word indices, n regional feature positions, and a regional visual dictionary capacity, where the size of the regional visual dictionary capacity is Num.

[0077] Among them, the regional visual word index can be expressed as The regional feature position can be expressed as The regional visual dictionary capacity can be expressed as Num. In this embodiment, the arrangement order and quantity of the regional visual word index and the regional feature position are defined to ensure the accuracy of the regional visual word bag, thereby further ensuring the accuracy of the hybrid visual word bag and subsequent regional similarity.

[0078] In one embodiment, the edge visual word bag includes m edge visual word indices, m edge feature positions, and an edge visual dictionary capacity, where the size of the edge visual dictionary capacity is Num.

[0079] Among them, the edge visual word index can be expressed as The edge feature position can be expressed as The edge visual dictionary capacity can be expressed as Num. In this embodiment, the arrangement order and quantity of the edge visual word index and the edge feature position are defined to ensure the accuracy of the edge visual word bag, thereby further ensuring the accuracy of the hybrid visual word bag and subsequent edge similarity.

[0080] By juxtaposing the regional visual word bag and the edge visual word bag to obtain the hybrid visual word bag, the information in the hybrid visual word bag is made more rich and accurate, so as to more precisely verify the initial image retrieval result in the subsequent post-processing program and obtain a more accurate image retrieval result.

[0081] In one embodiment, the process of feature quantization and index construction includes:

[0082] Clustering the region feature vectors into region visual words according to vector similarity;

[0083] Clustering the edge feature vectors into edge visual words according to vector similarity;

[0084] Merging the region visual words and the edge visual words to generate a visual dictionary;

[0085] Establishing a one-to-one correspondence between the image features of the image and the visual words in the visual dictionary to generate a visual word index, and establishing a one-to-one correspondence between the visual words in the visual dictionary and the image features of the image to generate an inverted index;

[0086] Locating the positions of the image features in the image to obtain feature positions, and using the visual word index, the feature positions, and the visual dictionary capacity to form a visual word bag.

[0087] In this embodiment, the purpose of clustering the regional feature vectors into regional visual words and the edge feature vectors into edge visual words according to vector similarity is that the number of image features is often in the thousands. In this step, clustering the image features with high similarity into the same visual words can reduce the computational and storage overhead in subsequent steps and improve the retrieval efficiency. The K-means clustering algorithm is used to cluster all the feature vectors. For example, there are a total of N vectors, which are clustered into K clusters. Each cluster surrounds a clustering center, and there are a total of K clustering centers. The vectors within the cluster have a high similarity, while the similarity between clusters is low. That is to say, all the feature vectors are used to generate a visual dictionary, and there are a total of K visual words in this visual dictionary. The capacity of the visual dictionary is K. Then, these visual words are used to represent the image: calculate the distances from the feature vectors in the image to these K visual words, and the visual word with the closest distance is the visual word corresponding to the feature. For example, a certain visual dictionary contains a total of 5 visual words: (word 1, word 2, word 3, word 4, word 5), and a certain image extracts 7 features: (feature 1, feature 2, feature 3, feature 4, feature 5, feature 6, feature 7). Map the features to the visual words with the closest distance: assume that features 1, 2, and 5 belong to word 1, features 3 and 7 belong to word 2, and features 4 and 6 belong to word 4. Then the visual word index of this image is (1, 1, 2, 4, 1, 4, 2). Define that in addition to including the index of the visual word corresponding to the feature in the visual word bag of an image, it also includes the position coordinates of the feature in this image and the capacity of the visual dictionary. We further assume that the positions of features 1-7 are a, b, c, d, e, f, g. Then the visual word bag of this image is: {(1, 1, 2, 4, 1, 4, 2), (a, b, c, d, e, f, g), 5}. Through this process, a reliable visual word bag can be obtained, ensuring that the visual words in the visual dictionary correspond to the image features.

[0088] In one of the embodiments, the vectorization process includes:

[0089] Extracting the visual word index of the visual word bag;

[0090] Arranging according to the frequency and order of appearance of the visual words in the visual word index;

[0091] Obtaining the visual word bag vector; the visual word bag vector includes the regional visual word bag vector and the edge visual word bag vector.

[0092] In this embodiment, by arranging according to the frequency and order of appearance of the visual words in the visual word index, a reliable visual word bag vector is obtained, ensuring the accuracy of information in subsequent similarity calculations.

[0093] In one of the embodiments, the feature matching process includes:

[0094] Calculate the cosine distance between the regional visual word bag vector of the query image and the regional visual word bag vectors of all reference images in the image database to obtain the regional cosine distance, and use 1 minus the regional cosine distance to obtain the regional similarity between the query image and all reference images;

[0095] Calculate the cosine distance between the edge visual word bag vector of the query image and the edge visual word bag vectors of all reference images in the image database to obtain the edge cosine distance, and use 1 minus the edge cosine distance to obtain the edge similarity between the query image and all reference images.

[0096] In this embodiment, the specific operation process is illustrated by way of example as follows:

[0097] The word bag vector of reference image 1 is (3, 2, 0, 2, 0);

[0098] The word bag vector of reference image 2 is (1, 1, 1, 0, 0);

[0099] The query image is (3, 2, 0, 2, 1);

[0100] Obviously, the cosine distance between the word bag vectors of reference image 1 and the query image is closer, so their similarity is higher. Therefore, reference image 1 is arranged in a more forward position in the initial image retrieval result, that is, arranged according to the size of the obtained similarity, and the larger the similarity, the more forward the arrangement, until the maximum capacity value of the initial image retrieval result is reached. Finally, the obtained initial image retrieval result is used as input and input into the subsequent post-processing program. This process reduces the number of image retrievals and excludes most of the results that do not match the query image.

[0101] In one of the embodiments, the post-processing process includes:

[0102] Extract the top K reference images with the highest to lowest hybrid similarity in the initial image retrieval result to obtain the first reference image set;

[0103] Use the feature positions and visual word indexes in the hybrid visual word bag to perform a secondary comparison between the reference images in the first reference image set and the query image;

[0104] Generate a new similarity index through the secondary comparison, and perform a descending order again according to the new similarity index to obtain the final image retrieval result.

[0105] In this embodiment, the first reference image set is obtained by extracting the top K reference images with the highest to lowest mixed similarity in the initial image retrieval results, which further reduces the workload during image retrieval and improves the retrieval efficiency. Moreover, it is worth noting that the information in the hybrid visual word bag is utilized to perform a secondary comparison between the reference images in the first reference image set and the query image, generating a new similarity metric. Since the information in the hybrid visual word bag is more abundant and accurate than that of a single-source visual word bag, the generated new similarity metric is more accurate, thereby further improving the retrieval accuracy.

[0106] It is worth noting that the pre-set region feature extraction algorithm and edge feature extraction algorithm are the KAZE algorithm. The conduction function used in the KAZE algorithm includes:

[0107]

[0108]

[0109]

[0110] where L σ is the Gaussian smoothed image, is the gradient of L σ and k is the diffusion control factor. g2 is the region feature extraction algorithm, which well preserves the image regions with larger widths; g3 is the edge feature extraction algorithm, which well preserves the boundary information of the image. Through this step of targeted feature extraction of the local part of the image, the extracted image features are more abundant and comprehensive.

[0111] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0112] Please refer to Figure 2 , in one embodiment, an image retrieval device 100 based on a hybrid KAZE algorithm for constructing a word bag model is provided, including a feature extraction module 11, a feature indexing module 12, a word bag mixing module 13, a feature matching module 14, a similarity mixing module 15, an initial arrangement module 16, and a post-processing module 17. Among them:

[0113] The feature extraction module 11 is used to extract the regional feature vector of the reference image by using a preset regional feature extraction algorithm, and extract the edge feature vector of the reference image by using a preset edge feature extraction algorithm; all the regional feature vectors of all the reference images in the image database are used to generate a regional visual dictionary, and all the edge feature vectors are used to generate an edge visual dictionary. The feature indexing module 12 is used to perform feature quantization and construct indexes on the regional feature vector and the edge feature vector respectively, obtain a regional visual word bag and an edge visual word bag, and vectorize the regional visual word bag and the edge visual word bag to obtain a regional visual word bag vector and an edge visual word bag vector; the visual word bag includes a visual word index, a feature position, and a visual dictionary capacity, and the visual word bag is a regional visual word bag or an edge visual word bag. The word bag mixing module 13 is used to perform a juxtaposition operation on the regional visual word bag and the edge visual word bag, obtain a mixed visual word bag, and store the mixed visual word bag in a post-processing program. The feature matching module 14 is used to perform feature matching on the regional visual word bag vector and the edge visual word bag vector respectively, obtain a regional similarity and an edge similarity. The similarity mixing module 15 is used to take the square root of the sum of the square of the edge similarity and the square of the regional similarity to obtain a mixed similarity. The initial sorting module 16 is used to sort the images in descending order of the mixed similarity to obtain an initial image retrieval result. The post-processing module 17 is used to perform post-processing based on the initial retrieval result in combination with the mixed visual word bag to obtain a final image retrieval result.

[0114] The above-mentioned image retrieval device 100 that constructs a word bag model based on the hybrid KAZE algorithm, through the cooperation of each module, uses a preset regional feature extraction algorithm to extract the regional feature vector of the reference image, and uses a preset edge feature extraction algorithm to extract the edge feature vector of the reference image. Then, the obtained regional feature vector and edge feature vector are subjected to feature quantization and index construction to obtain the corresponding visual word bag and visual word bag vector, and they are applied to the subsequent construction of the mixed word bag and the calculation of the mixed similarity, so that the mixed word bag can more completely describe the image, and the mixed similarity can more accurately reflect the similarity between the query image and the reference image. All the reference images are sorted in descending order of the mixed similarity to obtain an initial image retrieval result, and post-processing is performed based on the initial image retrieval result in combination with the mixed visual word bag to filter out unmatched images and reorder to obtain a final image retrieval result, thereby improving the overall accuracy of image retrieval and effectively improving the image retrieval accuracy.

[0115] In one embodiment, each module of the above-mentioned image retrieval device 100 that constructs a word bag model based on the hybrid KAZE algorithm can also be used to implement the corresponding processing steps of other embodiments of the above-mentioned image retrieval method that constructs a word bag model based on the hybrid KAZE algorithm.

[0116] For the specific limitations of the image retrieval device 100 based on the hybrid KAZE algorithm for constructing the bag-of-words model, reference can be made to the corresponding limitations of the image retrieval method based on the hybrid KAZE algorithm for constructing the bag-of-words model in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned image retrieval device 100 based on the hybrid KAZE algorithm for constructing the bag-of-words model can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of a device with specific data processing functions in the form of hardware, or stored in the memory of the foregoing device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules. The foregoing device can be, but is not limited to, various types of portable data analysis and processing devices existing in the art.

[0117] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following processing steps are implemented:

[0118] Extract the regional feature vector of the reference image by using a pre-set regional feature extraction algorithm, and extract the edge feature vector of the reference image by using a pre-set edge feature extraction algorithm; all the regional feature vectors of all the reference images in the image database are used to generate a regional visual dictionary, and all the edge feature vectors are used to generate an edge visual dictionary. Feature quantization and indexing are respectively performed on the regional feature vector and the edge feature vector to obtain a regional visual bag-of-words and an edge visual bag-of-words, and the regional visual bag-of-words and the edge visual bag-of-words are vectorized to obtain a regional visual bag-of-words vector and an edge visual bag-of-words vector; the visual bag-of-words includes a visual word index, a feature position, and a visual dictionary capacity, and the visual bag-of-words is a regional visual bag-of-words or an edge visual bag-of-words. A juxtaposition operation is performed on the regional visual bag-of-words and the edge visual bag-of-words to obtain a hybrid visual bag-of-words and store the hybrid visual bag-of-words in a post-processing program. Feature matching is respectively performed on the regional visual bag-of-words vector and the edge visual bag-of-words vector to obtain a regional similarity and an edge similarity. The square root of the sum of the square of the edge similarity and the square of the regional similarity is obtained as the hybrid similarity. The images are sorted in descending order according to the hybrid similarity to obtain an initial image retrieval result. Based on the initial retrieval result and in combination with the hybrid visual bag-of-words, post-processing is performed to obtain a final image retrieval result.

[0119] It can be understood that in addition to the above-mentioned memory and processor, the above-mentioned computer device further includes other software and hardware components not listed in this specification. Specifically, it can be determined according to the model of the specific data processing device in different application scenarios, and this specification will not list and elaborate one by one.

[0120] In one embodiment, when the processor executes the computer program, it can also implement the additional steps or sub-steps in each embodiment of the above-mentioned image retrieval method based on the hybrid KAZE algorithm for constructing the bag-of-words model.

[0121] In one embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the following processing steps are implemented:

[0122] Extract the region feature vector of the reference image by using a pre-set region feature extraction algorithm, and extract the edge feature vector of the reference image by using a pre-set edge feature extraction algorithm; all the region feature vectors of all the reference images in the image database are used to generate a region visual dictionary, and all the edge feature vectors are used to generate an edge visual dictionary. Feature quantization and indexing are respectively performed on the region feature vector and the edge feature vector to obtain a region visual word bag and an edge visual word bag, and the region visual word bag and the edge visual word bag are vectorized to obtain a region visual word bag vector and an edge visual word bag vector; the visual word bag includes a visual word index, a feature position, and a visual dictionary capacity, and the visual word bag is a region visual word bag or an edge visual word bag. A juxtaposition operation is performed on the region visual word bag and the edge visual word bag to obtain a mixed visual word bag and the mixed visual word bag is stored in a post-processing program. Feature matching is respectively performed on the region visual word bag vector and the edge visual word bag vector to obtain a region similarity and an edge similarity. The square root of the sum of the square of the edge similarity and the square of the region similarity is obtained as the mixed similarity. The images are sorted in descending order of the mixed similarity to obtain an initial image retrieval result. Post-processing is performed based on the initial retrieval result in combination with the mixed visual word bag to obtain a final image retrieval result.

[0123] In one embodiment, when the computer program is executed by a processor, the steps or sub-steps added in each of the above embodiments of the image retrieval method for constructing a word bag model based on the hybrid KAZE algorithm can also be implemented.

[0124] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus dynamic random access memory (Rambus DRAM, abbreviated as RDRAM), and interface dynamic random access memory (DRDRAM), etc.

[0125] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0126] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An image retrieval method for constructing a bag-of-words model based on a hybrid KAZE algorithm, characterized in that, The method includes: Extracting the region feature vector of the reference image by using a pre-set region feature extraction algorithm, and extracting the edge feature vector of the reference image by using a pre-set edge feature extraction algorithm; all the region feature vectors of all the reference images in the image database are used to generate a region visual dictionary, and all the edge feature vectors are used to generate an edge visual dictionary; Performing feature quantization and constructing an index on the region feature vector and the edge feature vector respectively to obtain a region visual word bag and an edge visual word bag, and vectorizing the region visual word bag and the edge visual word bag to obtain a region visual word bag vector and an edge visual word bag vector; the visual word bag includes a visual word index, a feature position, and a visual dictionary capacity, and the visual word bag is the region visual word bag or the edge visual word bag; Performing a juxtaposition operation on the region visual word bag and the edge visual word bag to obtain a mixed visual word bag and storing the mixed visual word bag in a post-processing program; Performing feature matching on the region visual word bag vector and the edge visual word bag vector respectively to obtain a region similarity and an edge similarity; Taking the square root of the sum of the square of the edge similarity and the square of the region similarity to obtain a mixed similarity; Sorting all the reference images in the image database in descending order according to the mixed similarity to obtain an initial image retrieval result; Performing post-processing based on the initial image retrieval result in combination with the mixed visual word bag to obtain a final image retrieval result; Among them, the process of the juxtaposition operation includes: Performing a juxtaposition operation on the visual word index in the region visual word bag and the visual word index in the edge visual word bag to form a mixed visual word index of the mixed visual word bag; Performing a juxtaposition operation on the feature position in the region visual word bag and the feature position in the edge visual word bag to form a mixed feature position of the mixed visual word bag; Performing a summation operation on the visual dictionary capacities in the region visual word bag and the edge visual word bag to form a mixed visual dictionary capacity; the mixed visual word index, the mixed feature position, and the mixed visual dictionary capacity form the mixed visual word bag; The process of the feature quantization and constructing an index includes: Clustering the region feature vector into region visual words according to vector similarity; Clustering the edge feature vector into edge visual words according to vector similarity; Merging the region visual words and the edge visual words to generate a visual dictionary; Generating a visual word index by corresponding the image features of the image with the visual words in the visual dictionary one by one, and generating an inverted index by corresponding the visual words in the visual dictionary with the image features of the image one by one; Locating the position of the image feature in the image to obtain the feature position, and forming the visual word bag by using the visual word index, the feature position, and the visual dictionary capacity; The process of the feature matching includes: Calculate the cosine distance between the regional visual word bag vector of the query image and the regional visual word bag vectors of all reference images in the image database to obtain the regional cosine distance, and use 1 minus the regional cosine distance to obtain the regional similarity between the query image and all reference images. The process of the post-processing includes: Extract the top K reference images with the highest to lowest mixed similarity in the initial image retrieval results to obtain the first set of reference images. Utilize the feature positions and visual word indexes in the mixed visual word bag to perform a secondary comparison between the reference images in the first set of reference images and the query image. Generate a new similarity metric through the secondary comparison, and perform the descending order according to the new similarity metric to obtain the final image retrieval result.

2. The image retrieval method for constructing a bag-of-words model based on a hybrid KAZE algorithm according to claim 1, characterized in that, The mixed visual word bag includes (n + m) mixed visual word indexes, (n + m) mixed feature positions, and a mixed visual dictionary capacity, where n is the number of regional features, m is the number of edge features, and the size of the mixed visual dictionary capacity is 2Num.

3. The image retrieval method for constructing a bag-of-words model based on a hybrid KAZE algorithm according to claim 1, characterized in that, The regional visual word bag includes n regional visual word indexes, n regional feature positions, and a regional visual dictionary capacity, where the size of the regional visual dictionary capacity is Num.

4. The image retrieval method for constructing a bag-of-words model based on a hybrid KAZE algorithm according to claim 1, characterized in that, The edge visual word bag includes m edge visual word indexes, m edge feature positions, and an edge visual dictionary capacity, where the size of the edge visual dictionary capacity is Num.

5. The image retrieval method for constructing a bag-of-words model based on a hybrid KAZE algorithm according to claim 1, characterized in that, The process of vectorization includes: Extract the visual word indexes of the visual word bag. Arrange according to the frequency and order of appearance of the visual words in the visual word indexes to obtain the visual word bag vector; the visual word bag vector includes the regional visual word bag vector and the edge visual word bag vector.

6. The image retrieval method for constructing a bag-of-words model based on the hybrid KAZE algorithm according to claim 1, wherein, The process of feature matching further includes: Calculate the cosine distance between the edge visual word bag vector of the query image and the edge visual word bag vectors of all reference images in the image database to obtain the edge cosine distance, and use 1 minus the edge cosine distance to obtain the edge similarity between the query image and all reference images.

Citation Information

Patent Citations

  • Image retrieval method and system

    CN108255858A

  • Coding and decoding method for images or videos

    US20150131921A1