Image retrieval method and device and readable storage medium
By combining HSV color histogram, color moment and deep neural network to extract features, the feature vectors to be queryed are generated, which solves the problems of low efficiency and insufficient accuracy of traditional fabric patterns, and achieves more efficient pattern matching.
Patent Information
- Application Number
- CN202510781027.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional fabric pattern search methods are inefficient and subjective. A single feature cannot fully reflect the visual characteristics of the fabric pattern, resulting in incorrect matching and inefficient search results.
Combining HSV color histogram, color moment and deep neural network to extract features, generate feature vectors to be queried through feature fusion, use feature vectors to be retrieved, and use feature vectors to search, and build inverted indexes for image matching.
It improves the accuracy and stability of fabric pattern retrieval, and can more accurately match complex and variable fabric patterns.
Smart Images

Figure CN120296195A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an image retrieval method, apparatus, and readable storage medium. Background Art
[0002] As a traditional pillar industry in China, the textile industry is undergoing a transformation from "manufacturing" to "intelligent manufacturing". In the context of digital transformation, the industry has a higher demand for the retrieval of fabric patterns.
[0003] Traditional fabric pattern retrieval mainly relies on manual experience comparison or simple computer analysis. This method is inefficient and highly subjective, and it is difficult to handle when the data volume is large. At present, content-based image retrieval technology has been gradually applied in the textile industry, but there are still many problems in the retrieval of complex and variable fabric patterns. For example, a single feature (such as texture) cannot fully reflect the visual characteristics of fabric patterns, which is prone to incorrect matching. For example, only using color features for retrieval without combining texture features, semantic features, etc., the retrieval results are not stable and accurate enough. Summary of the Invention
[0004] Based on this, it is necessary to provide an image retrieval method, apparatus, and readable storage medium for the above technical problems.
[0005] In a first aspect, an embodiment of this application provides an image retrieval method, and the method includes:
[0006] Extracting a first color feature based on the HSV color histogram of the image to be queried, and extracting a second color feature based on the color moment of the image to be queried;
[0007] Extracting the deep semantic feature of the image to be queried by using a deep neural network;
[0008] Performing feature fusion on the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried;
[0009] Retrieving a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of several query images.
[0010] In one of the embodiments, the extracting a first color feature based on the HSV color histogram of the image to be queried includes:
[0011] Dividing the H channel, S channel, and V channel of the image to be queried into multiple intervals respectively;
[0012] Counting the number of pixels of the image to be queried located in each channel interval respectively;
[0013] Based on the number of pixels of the image to be queried in each channel interval, an HSV histogram is obtained;
[0014] Based on the HSV color histogram, a first color feature is obtained.
[0015] In one embodiment, the extracting a second color feature based on the color moments of the image to be queried includes:
[0016] Calculate the color mean, color standard deviation, and color skewness of the H channel, S channel, and V channel of the image to be queried respectively;
[0017] Based on each of the color means, each of the color standard deviations, and each of the color skewnesses, color moments are obtained;
[0018] Based on the color moments, a second color feature is obtained.
[0019] In one embodiment, the extracting the deep semantic feature of the image to be queried by using a deep neural network includes:
[0020] Perform standard preprocessing on the image to be queried by using the image preprocessing function of the ResNet50 model;
[0021] Use the feature extraction layer of the ResNet50 model to extract the deep semantic feature of the preprocessed image to be queried, and obtain the deep semantic feature of the image to be queried.
[0022] In one embodiment, the fusing the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried includes:
[0023] Multiply the first color feature and the second color feature by a first weight respectively, and multiply the deep semantic feature by a second weight to obtain a feature vector to be queried, where the first weight is greater than the second weight.
[0024] In one embodiment, before retrieving the feature vector database based on the feature vector to be queried to obtain a retrieval result, the method further includes:
[0025] Divide the feature vectors of several query images in the feature vector database into multiple clusters, use each cluster center as an index entry, and construct an inverted index, where the inverted index records the feature vectors of the query images included in each cluster.
[0026] In one embodiment, the retrieving the feature vector database based on the feature vector to be queried to obtain a retrieval result includes:
[0027] Calculate the distances between the feature vector to be queried and all the cluster centers, and determine at least one target cluster;
[0028] Based on the inverted indexes of the respective target clusters, calculate the similarities between the feature vector to be queried and the feature vectors of the query images in the respective target clusters;
[0029] Based on the respective similarities, obtain the retrieval result.
[0030] In one embodiment, the calculating the similarities between the feature vector to be queried and the feature vectors of the query images in the respective target clusters based on the inverted lists of the respective target clusters includes:
[0031] Based on the inverted indexes of the respective target clusters, calculate the distances between the feature vector to be queried and the feature vectors of the query images in the respective target clusters;
[0032] Based on the respective distances, obtain the corresponding similarities.
[0033] In a second aspect, an embodiment of the present application further provides an image retrieval device, and the device includes:
[0034] A color feature extraction module, configured to extract a first color feature based on the HSV color histogram of the image to be queried, and extract a second color feature based on the color moments of the image to be queried;
[0035] A semantic feature extraction module, configured to extract the deep semantic feature of the image to be queried by using a deep neural network;
[0036] A feature vector construction module, configured to perform feature fusion on the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried;
[0037] A retrieval module, configured to retrieve a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of the feature vectors of a number of query images.
[0038] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, and a computer program is stored in the storage medium, wherein the computer program, when executed by a processor, implements the method described in the first aspect above.
[0039] The above image retrieval method, device, and readable storage medium extract a first color feature based on the HSV color histogram of the image to be queried, and extract a second color feature based on the color moments of the image to be queried; use a deep neural network to extract the deep semantic feature of the image to be queried; fuse the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried; retrieve a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of several query images, improving the accuracy of image retrieval.
[0040] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0042] Figure 1 is a hardware structure block diagram of a terminal device for an image retrieval method in an embodiment;
[0043] Figure 2 is a flowchart of an image retrieval method in an embodiment;
[0044] Figure 3 is an image to be queried in an embodiment;
[0045] Figure 4 is an HSV histogram in an embodiment;
[0046] Figure 5 is a structure block diagram of an image retrieval device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the purpose, technical solution, and advantages of the present application clearer and more understandable, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts fall within the scope of protection of the present application.
[0048] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0049] When "embodiment" is mentioned in the present application, it means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0050] Unless otherwise defined, the technical terms or scientific terms involved in the present application should have the ordinary meaning understood by those with ordinary skills in the technical field to which the present application belongs. The words such as "a", "an", "one kind", "the" and the like involved in the present application do not indicate a quantity limitation and can represent a single or plural number. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0051] The method embodiment provided in this embodiment can be executed on a terminal, a computer, or a similar computing device. For example, running on a terminal, Figure 1 is the hardware structure block diagram of the terminal of the image retrieval method in this embodiment. AsFigure 1 As shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0052] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the image retrieval method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0053] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0054] An embodiment of the present application provides an image retrieval method. Taking the terminal in Figure 1 as an example for illustration, as Figure 2 shown, the method includes the following steps:
[0055] Step 201, extract a first color feature based on the HSV color histogram of the image to be queried, and extract a second color feature based on the color moments of the image to be queried.
[0056] Step 202: Extract the deep semantic features of the image to be queried by using a deep neural network.
[0057] Step 203: Perform feature fusion on the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried.
[0058] Step 204: Retrieve a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of several query images.
[0059] In the above image retrieval method, the first color feature is extracted based on the HSV color histogram of the image to be queried, and the second color feature is extracted based on the color moment of the image to be queried; the deep semantic features of the image to be queried are extracted by using a deep neural network; the first color feature, the second color feature, and the deep semantic feature are subjected to feature fusion to obtain a feature vector to be queried; a feature vector database is retrieved based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of several query images, which improves the accuracy of image retrieval.
[0060] In one embodiment, the extracting the first color feature based on the HSV color histogram of the image to be queried includes the following:
[0061] Step 301: Divide the H channel, the S channel, and the V channel of the image to be queried into multiple intervals respectively.
[0062] The color histogram is the most commonly used color feature extraction method. By quantifying colors in the color space and counting the total number of pixel points in each quantized channel, the statistical distribution characteristics of the image colors are described. The advantage of the color histogram is that it can present the color composition in the form of an intuitive global histogram, effectively solving the color quantization problem under a complex texture background.
[0063] HSV is a color space that decomposes colors into three dimensions, where H (Hue) represents the hue, and its value range is [0, 180); S (Saturation) represents the saturation of the color, and its value range is [0, 255); V (Value) represents the brightness of the color, and its value range is [0, 255). In the embodiments of the present application, the H channel is evenly divided into 18 intervals, the S channel is evenly divided into 3 intervals, and the V channel is evenly divided into 3 intervals.
[0064] Step 302: Count the number of pixels of the image to be queried located in each channel interval respectively.
[0065] Step 303: Obtain an HSV histogram based on the number of pixels of the image to be queried located in each channel interval.
[0066] Step 304: Obtain the first color feature based on the HSV color histogram.
[0067] In this application, the cv2.calcHist function is used to calculate the three-dimensional histogram. For the HSV value (h, s, v) of each pixel, calculate the histogram interval it belongs to, and count the number of pixels in each three-dimensional interval (bin h , bin s , bin v ). Use formula (1) for statistics to form the histogram matrix. Normalize the histogram by cv2.normalize, and flatten the three-dimensional histogram (with a shape of 18×3×3) into a one-dimensional vector (162 dimensions), as shown in formula (2), to obtain the first color feature.
[0068] (1)
[0069] (2)
[0070] Where hist[i] represents the count value of the i-th interval in the HSV histogram. In the HSV histogram, the number of pixels in a specific combination interval of H, S, and V is the value of hist[i] corresponding to this position. is the histogram data after normalization.
[0071] Exemplarily, as Figure 3 shows a query image to be processed. The HSV histogram features of the query image are extracted using the above method. With the value range (H / S / V) of the corresponding channel as the abscissa and the number of pixels in each sub-interval as the ordinate, the resulting graph is the HSV histogram, as Figure 4 shows. After normalization, the feature values of each channel of H, S, and V are as follows:
[0072] HSV color histogram feature values:
[0073] H channel: 0.59773, 0.04375, 0.00233, 0.00048, 0.00025, 0.00004, 0.00000, 0.00000, 0.00000, 0.00000, 0.00090, 0.00080, 0.00000, 0.00000, 0.90000, 0.00000, 0.00001, 0.35540
[0074] S channel: 0.06055, 0.09611, 0.84334
[0075] V channel: 0.49761, 0.31869, 0.18370
[0076] If only the color histogram is used to represent the color features, there will be certain deficiencies in the retrieval of fabric patterns. The color histogram is insufficient in dealing with images containing texture and spatial information and is prone to losing some details in images with complex color distributions. Therefore, color moments need to be added to improve the extraction of color features.
[0077] Color moments are calculated based on the principle of moments to represent the color distribution of an image. It regards the color space of the image as a probability distribution and describes the statistical features of the color by calculating the moments of the color components. The advantage of color moments is that they can quickly reflect the overall characteristics of the color distribution and have a low feature dimension, making them very suitable for real-time retrieval scenarios.
[0078] In one embodiment, extracting the second color feature based on the color moments of the to-be-query image includes the following:
[0079] Step 401, calculate the color mean, color standard deviation, and color skewness of the H channel, S channel, and V channel of the to-be-query image respectively;
[0080] Step 402, obtain color moments based on each of the color means, each of the color standard deviations, and each of the color skewnesses; obtain the second color feature based on the color moments.
[0081] In the embodiments of the present application, three statistics, namely the mean (Mean) of the image color, the standard deviation (Std) of the image color, and the skewness (Skewness) of the image color, are calculated for each channel (H, S, V) of HSV respectively, integrated into a list, and finally converted into a 9-dimensional NumPy array. Among them, Mean can reflect the overall color level of the channel, and its calculation method is formula (3); Std can reflect the degree of dispersion of the image color distribution, and its calculation method is formula (4); Skewness can reflect the degree of asymmetry of its distribution, and its calculation method is formula (5).
[0082] (3)
[0083] (4)
[0084] (5)
[0085] Where N represents the total number of pixels in the color channel, x represents the specific value of a single pixel in the channel, channel represents one of H, S, and V, and sign(m3) represents taking the sign (positive or negative) of m3.
[0086] Exemplarily, see Figure 3For the displayed image to be queried, three statistical measures (mean, standard deviation, skewness) are calculated for each channel using the above method. The 9-dimensional feature vector obtained from the three HSV channels is the color moment, as follows:
[0087] Color moment eigenvalues: [64.681046 84.368095 71.16232
[0089] 214.5537 55.395695 -65.21314 101.70508 70.24241 58.59078]
[0091] In one embodiment, the extraction of the deep semantic features of the image to be queried using a deep neural network includes the following: performing standardization preprocessing on the image to be queried using the image preprocessing function of the ResNet50 model; using the feature extraction layer of the ResNet50 model to extract deep semantic features from the preprocessed image to be queried, obtaining the deep semantic features of the image to be queried.
[0092] Using only color features as the sole feature for fabric pattern retrieval, while lightweight and efficient, has limited semantic coverage and is only suitable for simple color matching scenarios. To address this need, deep semantic features are selected to be added to enhance semantic understanding, enabling the system to achieve accurate retrieval even when faced with complex patterns.
[0093] Comparing with other models such as VGG, Inception, MobileNet, ResNet101 / 152, etc., ResNet50 balances accuracy and computational complexity, having a faster inference speed while maintaining high accuracy. The core of ResNet50 is residual learning, which solves the gradient vanishing problem of deep neural networks by introducing residual blocks. The deep convolutional layers of ResNet50 can capture abstract semantic information such as pattern structures and object contours, making up for the deficiencies of color features.
[0094] First, in the part of loading the pre-trained ResNet50 model, remove the last classification layer and retain the feature extraction layer, setting the output feature dimension to N dimensions; then start preprocessing the image. The preprocessing uses the algorithm provided by the model, obtaining the recommended image preprocessing function from the model and performing standardization preprocessing; then extract the deep semantic features of the image. See formula (6), the model outputs a tensor with a shape of [1, N, 1, 1], which is compressed to obtain an N-dimensional feature vector. Exemplarily, N is set to 2048.
[0095] (6)
[0096] Where It is the feature extraction part of the ResNet50 model (removing the classification layer); img_tensor is the preprocessed image tensor with a shape of [1, 3, 224, 224]; features is the output N-dimensional feature vector representing the deep semantic information of the image.
[0097] Taking Figure 3 the query image shown as an example, use the ResNet50 pre-trained model to calculate its deep semantic features, obtaining a 2048-dimensional feature vector data. The first 20 and the last 20 eigenvalue are shown in Table 1. At the same time, to ensure the accuracy and readability of the data, all the displayed data are retained to five decimal places.
[0098] Table 1
[0099]
[0100] In one embodiment, the step of fusing the first color feature, the second color feature, and the deep semantic feature to obtain a query feature vector includes the following: multiplying the first color feature and the second color feature by a first weight respectively, and multiplying the deep semantic feature by a second weight to obtain the query feature vector.
[0101] Since a single feature often has limitations, in order to describe the image content more comprehensively and accurately, this application selects to fuse color features and deep semantic features. Compared with simply concatenating the 171-dimensional color feature (the first color feature and the second color feature) and the 2048-dimensional deep semantic feature into a 2219-dimensional vector, this application assigns certain weights to the color feature and the deep feature respectively. By adjusting the proportion of different features through weights, it can avoid a single feature dominating the retrieval result and balance the feature scale difference.
[0102] In one exemplary embodiment, for the same image, in each round of testing, retrieval experiments are sequentially carried out on a variety of different weight combination ratios. Finally, the first color feature and the second color feature are given an 80% weight, and the deep semantic feature is given a 20% weight. It can be understood that the weight ratio of the first color feature and the second color feature is greater than that of the deep semantic feature, and the retrieval accuracy is higher.
[0103] In one embodiment, before retrieving the feature vector database based on the query feature vector to obtain a retrieval result, the method further includes: dividing the feature vectors of several query images in the feature vector database into multiple clusters, using each cluster center as an index entry to construct an inverted index, and the inverted index records the feature vectors of the query images included in each cluster.
[0104] Specifically, this application uses the above method for constructing the feature vector of the image to be queried to build a feature vector database based on the query image dataset. If each feature vector in the feature vector database is directly traversed, a large amount of memory is occupied when calculating the distance between the feature vector to be queried and all vectors in the database every time a retrieval is performed. Therefore, this application adds a Faiss index to the database. Compared with the original system traversing the database, the Faiss index can be constructed in advance and directly used during each search, which can reduce the computational amount required for querying.
[0105] The process of index construction is as follows: First, the function divides the feature vectors into nlist clusters, and each cluster corresponds to a center point. See formula (7). The process of calculating the cluster center is called training the index. After training the index, the feature vectors need to be assigned to the nearest cluster.
[0106] (7)
[0107] Its optimization goal is to minimize the sum of the squares of the distances from all data points to the center of the cluster they belong to. Among them, nlist is the number of clusters, Ci represents the i-th cluster, x j is the j-th data point belonging to Ci, and µ i is the center (mean vector) of the i-th cluster, and ‖x - µi‖ 2 refers to the square of the Euclidean distance from the data point x to the center of the cluster µ i it belongs to.
[0108] In one embodiment, retrieving the feature vector database based on the feature vector to be queried to obtain a retrieval result includes: calculating the distances between the feature vector to be queried and all cluster centers, and determining at least one target cluster; calculating the similarities between the feature vector to be queried and the feature vectors of each query image in each target cluster based on the inverted index of each target cluster; and obtaining a retrieval result based on each similarity.
[0109] In one embodiment, calculating the similarities between the feature vector to be queried and the feature vectors of each query image in each target cluster based on the inverted list of each target cluster includes: calculating the distances between the feature vector to be queried and the feature vectors of each query image in each target cluster based on the inverted index of each target cluster; and obtaining the corresponding similarities based on each distance.
[0110] In the retrieval stage, the inverted index of Faiss is mainly used to accelerate the search. First, it is necessary to calculate the distance between the query vector q and all cluster centers Ci, and select the nearest nprobe clusters, as shown in formula (8). In the vector list of the selected clusters, calculate the L2 distance between the query vector and each candidate vector, as shown in formula (9). Finally, return the K nearest neighbor vectors.
[0111] (8)
[0112] (9)
[0113] Exemplarily, the default value of nprobe is 1, distance(q, x) represents the squared Euclidean distance between the query vector q and the candidate vector x, d represents the dimension of the vector, that is, the number of features of each query vector q and candidate vector x (2219 dimensions), and the value range of i is from 1 to d.
[0114] Map the calculated original distance data to the range of 0 to 1, that is, the required similarity score, as shown in formula (10). Finally, sort the retrieved image paths and similarity scores in descending order of similarity and output, and return the K most similar vectors, thus completing the similarity matching based on the Faiss index.
[0115] (10)
[0116] Among them, similarity is the required similarity score, distance is the distance value calculated between the query vector and a certain candidate vector, max_distance is the maximum value of all distance values obtained in this retrieval (that is, all distances distance), and the number of distances in this application is 4.
[0117] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.
[0118] In one embodiment, as Figure 5 shown, the embodiment of the present application also provides an image retrieval device, and the device includes:
[0119] A color feature extraction module 10, configured to extract a first color feature based on the HSV color histogram of the query image, and extract a second color feature based on the color moment of the query image;
[0120] A semantic feature extraction module 20, configured to extract deep semantic features of the image to be queried by using a deep neural network;
[0121] A feature vector construction module 30, configured to perform feature fusion on the first color feature, the second color feature, and the deep semantic features to obtain a feature vector to be queried;
[0122] A retrieval module 40, configured to retrieve a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of a plurality of query images.
[0123] In one embodiment, the color feature extraction module 10 is further configured to: divide the H channel, the S channel, and the V channel of the image to be queried into multiple intervals respectively; count the number of pixels of the image to be queried located in each channel interval respectively; obtain an HSV histogram based on the number of pixels of the image to be queried located in each channel interval; and obtain a first color feature based on the HSV color histogram.
[0124] In one embodiment, the color feature extraction module 10 is further configured to: calculate the color mean, the color standard deviation, and the color skewness of the H channel, the S channel, and the V channel of the image to be queried respectively; obtain color moments based on each of the color means, each of the color standard deviations, and each of the color skewnesses; and obtain a second color feature based on the color moments.
[0125] In one embodiment, the semantic feature extraction module 20 is further configured to: perform standardization preprocessing on the image to be queried by using an image preprocessing function of a ResNet50 model; and perform deep semantic feature extraction on the preprocessed image to be queried by using a feature extraction layer of the ResNet50 model to obtain deep semantic features of the image to be queried.
[0126] In one embodiment, the feature vector construction module 30 is further configured to: multiply the first color feature and the second color feature by a first weight respectively, and multiply the deep semantic feature by a second weight to obtain a feature vector to be queried, where the first weight is greater than the second weight.
[0127] In one embodiment, the retrieval module 40 is further configured to: divide the feature vectors of a plurality of query images in the feature vector database into multiple clusters, use each cluster center as an index entry, and construct an inverted index, where the inverted index records the feature vectors of the query images included in each cluster.
[0128] In one embodiment, the retrieval module 40 is further configured to: calculate the distances between the feature vector to be queried and all the cluster centers, and determine at least one target cluster; calculate the similarities between the feature vector to be queried and the feature vectors of the query images in each of the target clusters based on the inverted index of each of the target clusters; and obtain a retrieval result based on each of the similarities.
[0129] In one embodiment, the retrieval module 40 is further configured to: calculate the distances between the feature vector to be queried and the feature vectors of the query images in each of the target clusters based on the inverted index of each of the target clusters; and obtain the corresponding similarities based on each of the distances.
[0130] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any of the above-described embodiments of the image retrieval method are implemented.
[0131] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0132] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0133] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An image retrieval method, characterized in that, The method includes: extracting a first color feature based on the HSV color histogram of the image to be queried, and extracting a second color feature based on the color moments of the image to be queried; extracting the deep semantic feature of the image to be queried by using a deep neural network; performing feature fusion on the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried; retrieving a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of a plurality of query images.
2. The method according to claim 1, characterized in that The extracting a first color feature based on the HSV color histogram of the image to be queried includes: dividing the H channel, the S channel, and the V channel of the image to be queried into multiple intervals respectively; counting the number of pixels of the image to be queried in each channel interval respectively; obtaining an HSV histogram based on the number of pixels of the image to be queried in each channel interval; obtaining a first color feature based on the HSV color histogram.
3. The method according to claim 1, wherein The extracting a second color feature based on the color moments of the image to be queried includes: calculating the color mean, the color standard deviation, and the color skewness of the H channel, the S channel, and the V channel of the image to be queried respectively; obtaining color moments based on each of the color means, each of the color standard deviations, and each of the color skewnesses; obtaining a second color feature based on the color moments.
4. The method according to claim 1, wherein The extracting the deep semantic feature of the image to be queried by using a deep neural network includes: performing standard preprocessing on the image to be queried by using the image preprocessing function of the ResNet50 model; extracting the deep semantic feature of the preprocessed image to be queried by using the feature extraction layer of the ResNet50 model to obtain the deep semantic feature of the image to be queried.
5. The method according to claim 1, characterized in that, The performing feature fusion on the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried includes: multiplying the first color feature and the second color feature by a first weight respectively, and multiplying the deep semantic feature by a second weight to obtain a feature vector to be queried, where the first weight is greater than the second weight.
6. The method according to claim 1, characterized in that, Before retrieving the feature vector database based on the feature vector to be queried to obtain a retrieval result, the method further includes: dividing the feature vectors of a plurality of query images in the feature vector database into multiple clusters, taking each cluster center as an index entry, and constructing an inverted index, where the inverted index records the feature vectors of the query images included in each cluster.
7. The method according to claim 6, characterized in that, The retrieving the feature vector database based on the feature vector to be queried to obtain a retrieval result includes: calculating the distances between the feature vector to be queried and all cluster centers, and determining at least one target cluster; calculating the similarities between the feature vector to be queried and the feature vectors of the query images in each of the target clusters based on the inverted index of each of the target clusters; obtaining a retrieval result based on each of the similarities.
8. The method according to claim 7, wherein The calculating the similarities between the feature vector to be queried and the feature vectors of the query images in each of the target clusters based on the inverted list of each of the target clusters includes: Based on the inverted index of each of the target clusters, calculate the distances between the feature vector to be queried and the feature vectors of each query image in each of the target clusters; Based on each of the distances, obtain the corresponding similarity.
9. An image retrieval device, characterized in that, The device includes: A color feature extraction module, configured to extract a first color feature based on the HSV color histogram of the image to be queried, and extract a second color feature based on the color moments of the image to be queried; A semantic feature extraction module, configured to extract the deep semantic feature of the image to be queried by using a deep neural network; A feature vector construction module, configured to perform feature fusion on the first color feature, the second color feature, and the deep semantic feature to obtain a feature vector to be queried; A retrieval module, configured to retrieve a feature vector database based on the feature vector to be queried to obtain a retrieval result; the feature vector database is composed of feature vectors of a plurality of query images.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method for extracting landmark scene abstract
CN101777059A
A method and apparatus for image retrieval
CN109033308A
Image retrieval method and system based on high-level semantic features and color features
CN110162657A
Garment image retrieval method fusing color feature and residual network depth feature
CN110825899A
Image retrieval system and method based on color moment and deep learning features
CN119493875A