Image search optimization method and device, equipment and storage medium

By performing structured partitioning and optimization on the image feature set, the challenges of retrieval accuracy and efficiency in image search technology are solved, and the retrieval performance of large image databases is improved.

CN120744157APending Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510805637.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing image search technologies face challenges in retrieval accuracy and efficiency in large image databases, especially when performing nearest neighbor searches on massive vectors in high-dimensional spaces, where the computational overhead and time delay are unacceptable.

Method used

By performing structured partitioning and optimization on the image feature set, including clustering, representative feature extraction and feature compression, an optimized feature set is formed to improve the quality of the feature set, thereby enhancing the retrieval accuracy and efficiency.

Benefits of technology

It significantly improves the accuracy and efficiency of image retrieval, reduces redundant information, reduces storage requirements and the computational complexity of online retrieval, and achieves a balance between accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744157A_ABST
    Figure CN120744157A_ABST
Patent Text Reader

Abstract

The invention provides an image search optimization method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large models and image search. According to the specific implementation scheme, a plurality of image features of a plurality of categories are acquired from a first feature set corresponding to an image base library; dividing the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space; according to the feature distribution characteristics of the plurality of feature intervals, performing optimization processing on at least one feature interval to obtain an optimized second feature set; and based on the second feature set, retrieving and classifying the target query image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of computer vision, deep learning, large models, and image search technology. Background Art

[0002] Image search and retrieval is a technology that has developed rapidly in recent years. It allows users to use an image as a query input to find and return several images with the most similar content in a large image database. Summary of the Invention

[0003] The present disclosure provides an image search method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided an image search optimization method, comprising:

[0005] Acquire multiple image features of multiple categories from a first feature set corresponding to the image base library;

[0006] Based on the distribution of the multiple image features in the feature space, the multiple image features are divided into multiple feature intervals;

[0007] Optimizing at least one feature interval according to feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set; and

[0008] Based on the second feature set, the target query image is retrieved and classified.

[0009] According to another aspect of the present disclosure, a graph search and retrieval optimization device is provided, comprising:

[0010] A feature acquisition module is used to acquire multiple image features of multiple categories from a first feature set corresponding to the image base library;

[0011] An interval division module, configured to divide the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space;

[0012] an optimization module, configured to optimize at least one feature interval according to feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set; and

[0013] The retrieval module is used to retrieve and classify the target query image based on the second feature set.

[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0020] The technical solutions of the embodiments of the present disclosure can specifically optimize the internal structure of the feature set to collaboratively improve the accuracy and efficiency of retrieval.

[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0023] Figure 1 is a flowchart of an image search optimization method according to an embodiment of the present disclosure;

[0024] Figure 2 is a flowchart of a specific embodiment of an image search optimization method according to an embodiment of the present disclosure;

[0025] Figure 3 is a flowchart of a specific embodiment of an image search optimization method according to an embodiment of the present disclosure;

[0026] Figure 4 is a structural diagram of an image search optimization method according to an embodiment of the present disclosure;

[0027] Figure 5 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] In related technologies, the implementation process of image search technology usually includes two main stages: offline feature extraction and online retrieval. In the offline stage, the system uses a deep learning model (such as a convolutional neural network CNN) as an image encoder to convert each original image in the database (also known as the "base library") into a high-dimensional, compact feature vector. All these feature vectors together constitute a feature set for retrieval. In the online stage, the system performs the same feature extraction process on the query image input by the user, and then uses the query feature vector as a benchmark to measure the similarity by calculating the distance between vectors in the feature set (such as cosine distance or Euclidean distance), and returns several results with the closest distance.

[0030] To improve the accuracy of image search and retrieval processes, the development of related technologies has primarily focused on optimizing the discriminative power of feature vectors. A mainstream approach is to start with model training, designing more sophisticated loss functions to guide the optimization of deep learning models. This approach aims to ensure that the features learned by the model have better distribution characteristics in the feature space. Specifically, image features of the same category are clustered more closely together, while features of different categories are pushed further apart.

[0031] On the other hand, as image databases grow ever larger (reaching millions, tens of millions, or even hundreds of millions), retrieval efficiency becomes another key challenge. Performing precise, exhaustive nearest neighbor searches on massive vectors in high-dimensional space requires unacceptable computational overhead and latency.

[0032] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the embodiments of the present disclosure provide an image search optimization method. By utilizing the technical solutions of the embodiments of the present disclosure, the internal structure of the feature set can be optimized in a targeted manner to collaboratively improve the accuracy and efficiency of retrieval.

[0033] Figure 1 FIG. 1 is a flow chart of an image search optimization method provided by an embodiment of the present disclosure. Figure 1 As shown, the method comprises at least the following steps:

[0034] S110 , obtaining multiple image features of multiple categories from a first feature set corresponding to an image database.

[0035] In the disclosed embodiment, an image base to be optimized is first obtained. The image base can be understood as a collection of original images, such as a database containing tens of thousands of different product images. For this image base, a pre-trained deep learning model (such as ResNet, ViT, etc.) is used to extract the features of each image, thereby obtaining a first feature set. Image features are high-dimensional vectors, such as a 512-dimensional floating-point vector, which describe the visual content of the image through vectors.

[0036] S120 : Divide the multiple image features into multiple feature intervals based on the distribution of the multiple image features in the feature space.

[0037] The feature space is a multi-dimensional mathematical space, and each image feature is a point in this space. This step aims to analyze the distribution of features of the same category in space, such as which features are concentrated in the center and which are distributed at the edge. Through specific division rules (which will be described in detail in Example 2), the features of the same category are divided into different feature intervals, that is, the features within a category are structured and grouped according to their distribution characteristics. Each category is associated with multiple feature intervals. Each feature interval represents a different area in the feature space.

[0038] S130 . Optimize at least one feature interval according to feature distribution characteristics of the multiple feature intervals to obtain an optimized second feature set.

[0039] Based on the feature distribution characteristics of each feature interval (e.g., the density or representativeness of the interval), at least one feature interval is optimized. This optimization can be performed to improve retrieval accuracy or efficiency. Optimization methods include, but are not limited to, generating representative features for the interval to improve accuracy or compressing features within the interval to improve efficiency. After processing, the original first feature set is converted or enhanced into a second feature set.

[0040] S140 : Retrieve and classify the target query image based on the second feature set.

[0041] When a target query image to be retrieved is obtained, the optimized second feature set is used as the retrieval object, and similarity calculation and classification judgment are performed to return the final retrieval result.

[0042] According to the solution of the embodiment of the present disclosure, by performing structured partitioning and optimization on the feature set based on its intrinsic distribution, the quality of the feature set can be improved in a targeted manner, thereby achieving subsequent retrieval accuracy and efficiency.

[0043] In one possible implementation, S120 divides the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space, further including:

[0044] S121 . For each target category in the multiple categories, cluster multiple image features included in the target category to obtain a cluster center.

[0045] S122 . Sort the multiple image features according to the distances from the multiple image features included in the target category to the cluster center to obtain a sorting result.

[0046] S123. Divide the multiple image features into multiple ordered feature intervals according to the sorting result.

[0047] In this disclosed embodiment, a clustering algorithm, such as the K-Means algorithm (with K set to 1), is applied to all image features of a target category (e.g., "cat"). The goal of clustering is to find the centroid of this group of features in the feature space. This centroid is the cluster center, which numerically represents the most typical and core average feature of the category.

[0048] Calculate the distance between each image feature vector in that category and the cluster center vector obtained in the previous step. Use either Euclidean distance or cosine distance as the distance metric. Sort all image features in ascending order based on the calculated distances, resulting in a ranking from most typical to least typical.

[0049] Based on the sorting results of the previous step, in one example, a preset distance percentage step (e.g., 10%) can be used to segment the features into multiple ordered intervals. Specifically, the features that are 0% to 10% of the first sort are assigned to the first interval, the features that are 10% to 20% of the first sort are assigned to the second interval, and so on, until the features that are 90% to 100% of the first sort are assigned to the tenth interval.

[0050] In another example, the feature intervals can be divided into fixed numbers. The specific implementation process is as follows:

[0051] Set a preset interval size N, for example, N = 1000. This value can be selected based on experience or according to the granularity requirements of subsequent optimization tasks.

[0052] Starting from the beginning of the sorted list (that is, the feature closest to the cluster center), perform the following operations in sequence:

[0053] Take the 1st to Nth features in the list and group them into "Feature Interval One".

[0054] Next, take the features from N+1th to 2Nth in the list and classify them into "Feature Interval 2".

[0055] This process is repeated until all features in the list have been assigned.

[0056] If the total number of features is not an integer multiple of N, then after the last complete interval is allocated, there will be some remaining features at the end of the list. These remaining features that are less than N will form the last feature interval.

[0057] According to the solution of the embodiment of the present disclosure, the fuzzy feature distribution within a category can be converted into structured and quantifiable multiple levels (feature intervals), providing an accurate operational basis for subsequent targeted optimization.

[0058] In one possible implementation, S130 optimizes at least one feature interval according to feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set, including:

[0059] S131 . For each of the multiple feature intervals, calculate the mean of all image features in the feature interval to obtain representative features of the multiple feature intervals.

[0060] S132. Add at least one representative feature to the first feature set to obtain an optimized second feature set.

[0061] In the disclosed embodiment, a representative feature is calculated for each defined feature interval. Specifically, this calculation may be performed by taking the dimension-by-dimension arithmetic average of all image feature vectors within that interval. The resulting mean vector is the representative feature for that interval. For example, the representative feature for the first feature interval summarizes the average appearance of the most core 10% of images in that category.

[0062] One or more representative features calculated in the previous step are directly added back to the original first feature set to form a second feature set that is larger but contains more robust information.

[0063] According to the solution of the embodiment of the present disclosure, the representative features are like adding several standard images with strong anti-interference capabilities to each category. During retrieval, the query image can be matched with these more general features, thereby effectively improving the accuracy and robustness of the retrieval.

[0064] In a possible implementation, S132 adds at least one representative feature to the first feature set to obtain an optimized second feature set, which may be:

[0065] S132-1. Add representative features of at least one feature interval among the multiple feature intervals to the first feature set to obtain an optimized second feature set.

[0066] In the disclosed embodiment, some categories may only require a representative feature in their most core feature interval (such as the 0% to 10% interval) to achieve optimal accuracy improvement. In this case, only this one representative feature may be added.

[0067] In another possible implementation, S132 adds at least one representative feature to the first feature set to obtain an optimized second feature set, which may be:

[0068] S132-2. Arrange and combine representative features of at least some of the feature intervals in the plurality of feature intervals and add them to the first feature set to obtain an optimized second feature set.

[0069] In the disclosed embodiment, for other categories with richer visual features, it may be necessary to combine the representative features of the core interval (such as 0% to 10%) and a certain intermediate interval (such as 40% to 50%) to form a combination and add them together to the feature set to cover the various main appearances of the category.

[0070] According to the solution of the embodiment of the present disclosure, a variety of joining strategies are provided, so that the optimization solution can be adaptively adjusted according to the different characteristics of each category, realizing "one category, one strategy" and maximizing the accuracy improvement effect.

[0071] In one possible implementation, Figure 2 As shown, after executing S120, S130 optimizes at least one feature interval according to the feature distribution characteristics of the multiple feature intervals to obtain an optimized second feature set, specifically including:

[0072] S210: Count the number of image features in multiple feature intervals.

[0073] S220 . For a dense feature interval in which the number of image features exceeds a first threshold, compress the image features in the dense feature interval in terms of quantity at a first compression rate to obtain an optimized second feature set.

[0074] In the embodiment of the present disclosure, the feature intervals divided previously can be used, or steps S120 or S121 to S123 can be repeated to divide new feature intervals using a different preset distance percentage step size (for example, a step size of 20%, divided into 5 feature intervals), and the image features in each interval can be counted.

[0075] A dense feature interval is an interval with a very high number of features, typically near the cluster center. If the number of features in a feature interval exceeds a first threshold (e.g., 15% of the total number of features in a category), it is compressed. For example, random sampling is performed at a first compression rate (e.g., 50%), retaining only 50% of the features. After compression, part of the first feature set is removed, resulting in a smaller second feature set.

[0076] The solution of the embodiment of the present disclosure can significantly reduce redundant information in the feature set because the similarity between features in dense feature intervals is extremely high. Through compression, the storage requirements and the computational complexity of online retrieval can be greatly reduced, thereby significantly improving retrieval efficiency.

[0077] In a possible implementation, based on S120, S210, and S220, the image search optimization method of the embodiment of the present disclosure further includes the following steps:

[0078] S230: For the sparse feature intervals where the number of image features is not greater than the second threshold, do not compress the number of image features or compress the number of image features at a second compression rate to obtain an optimized second feature set, where the second compression rate is lower than the first compression rate.

[0079] In the disclosed embodiment, a sparse feature interval generally refers to an interval far from the cluster center, which has a small number of features. To this end, a second threshold is set (which may be equal to or less than the first threshold). For intervals where the number of features does not exceed this threshold, it can be considered that the features contained therein are relatively unique and have high information value, so no compression is performed, or only a slight compression is performed at a very low second compression rate (such as 10%, which is less than the first compression rate of 50%).

[0080] According to the solution of the embodiment of the present disclosure, by protecting sparse intervals, it is possible to effectively avoid excessive loss of retrieval recall rate while pursuing efficiency, retain the ability to recognize images that are non-mainstream in the category but have important distinguishing significance, and achieve a better balance between accuracy and efficiency.

[0081] In a possible implementation, S140 retrieves and classifies the target query image based on the second feature set, further comprising the steps of:

[0082] S141. Determine query features from the target query image.

[0083] This step is the same as the feature extraction in the offline stage, that is, the target query image uploaded by the user is converted into a query feature vector using the same image encoder.

[0084] S142: Determine similarity scores between the query feature and multiple features in the second feature set.

[0085] Calculate the similarity between the query feature and each feature vector in the second feature set (the feature set optimized by any of the above methods), for example, calculate the cosine similarity, and obtain a series of similarity scores and corresponding category labels.

[0086] S143: Determine the category to which the target query image belongs based on a preset voting strategy and similarity scores.

[0087] In the disclosed embodiments, the preset voting strategy is a decision rule used to comprehensively determine the most reliable category assignment from a large number of similarity scores. This strategy can be as simple as taking the one with the highest similarity or a more complex statistical algorithm (detailed in the embodiments below).

[0088] According to the solution of the embodiment of the present disclosure, a complete search request is completed using the optimized feature set, thereby improving the search efficiency and search accuracy.

[0089] In one possible implementation, Figure 3 As shown, the image search optimization method of the embodiment of the present disclosure further includes the steps of:

[0090] S310 , for multiple categories in the image database, based on historical retrieval performance data or feature distribution data of the categories, select an optimal voting strategy for the category from multiple candidate voting strategies as a preset voting strategy.

[0091] The disclosed embodiment provides an offline configuration optimization process. First, prepare an independent, labeled test set. For each category in the base library (such as "cat"), use all candidate voting strategies (such as Softmax strategy, threshold voting strategy, etc.) to run a complete retrieval test on the test set in turn. Then, evaluate the performance of each strategy on the category based on historical retrieval performance data (such as F1-score, Precision, Recall and other indicators). For example, if the "threshold voting strategy" obtains the highest F1-score on the "cat" category, the system will record the "threshold voting strategy" as the optimal voting strategy for the "cat" category. After this process is executed for all categories, a set of "one category, one strategy" optimal voting strategy configuration tables is formed.

[0092] According to the solution of the embodiment of the present disclosure, the voting decision can match the most suitable decision algorithm according to the data characteristics of each category, thereby maximizing the final accuracy of the retrieval.

[0093] In a possible implementation, S143 determines the category to which the target query image belongs based on a preset voting strategy and a similarity score, further comprising the steps of:

[0094] The similarity scores are normalized using the soft maximum function to obtain normalized scores.

[0095] The category with the highest normalized score is determined as the category to which the target query image belongs.

[0096] In the disclosed embodiment, the Softmax function is a function that converts any real number vector into a probability distribution. All similarity scores obtained in S142 are input to the Softmax function. The sum of the normalized scores output by the Softmax function is 1, which can be regarded as the probability that the query image belongs to each candidate category. The category with the largest normalized score is directly selected as the final decision result.

[0097] According to the solution of the embodiment of the present disclosure, the Softmax strategy is simple and has a fast calculation speed, and is suitable for scenarios where the distinction between category features is very obvious.

[0098] In a possible implementation, S143 determines the category to which the target query image belongs based on a preset voting strategy and a similarity score, further comprising the steps of:

[0099] The similarity score is compared with a preset similarity threshold.

[0100] For each category, count the number of times the similarity score of its features exceeds the similarity threshold to get a vote.

[0101] The category with the highest number of votes is determined as the category to which the target query image belongs.

[0102] In the disclosed embodiment, a global similarity threshold is set, for example, 0.85. All similarity scores obtained in S143 are traversed. If a score is greater than 0.85, the corresponding category receives one vote. After this process is complete, each category receives a final number of votes. The category with the most votes is selected as the final search result.

[0103] According to the solution of the embodiment of the present disclosure, this strategy is insensitive to high scores of outliers, and pays more attention to the number of features similar to the query image. It has good robustness and is suitable for scenarios with diverse morphologies within the category.

[0104] In a possible implementation, S143 determines the category to which the target query image belongs based on a preset voting strategy and a similarity score, further comprising the steps of:

[0105] The similarity scores are sorted in descending order, and the top K similarity scores and their corresponding candidate categories are selected, where K is a preset integer greater than 1.

[0106] For each candidate category, the similarity scores of all the top K selected categories are accumulated to obtain a total category score.

[0107] The category with the highest total category score is determined as the category to which the target query image belongs.

[0108] The value of K is 3, 5 or 9.

[0109] In the disclosed embodiment, all similarity scores obtained in S142 are first sorted from high to low. Then, the top K results in the sorting results are selected. K is a preset integer, for example, it can be set to 3, 5 or 9 based on experience or experiments. The selected Top-K results may include multiple different categories. This step adds up all the scores belonging to the same candidate category to obtain the total category score of the category. For example, in the Top-5 results, "cat" appears 3 times and "dog" appears 2 times, then the scores of the 3 "cats" are added together, and the scores of the 2 "dogs" are added together. Finally, the total category scores of all candidate categories are compared, and the category with the highest score is the final classification result.

[0110] According to the solution of the embodiment of the present disclosure, this strategy comprehensively considers the top-ranked results, taking into account both the similarity level and the number of high-similarity results, and is a voting strategy with generally excellent overall performance.

[0111] Figure 4 FIG. 4 is a structural diagram of an image search optimization device 400 according to an embodiment of the present disclosure. Figure 4 As shown, the device includes:

[0112] The feature acquisition module 401 is configured to acquire a plurality of image features of a plurality of categories from a first feature set corresponding to an image base database.

[0113] The interval division module 402 is configured to divide the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space.

[0114] The optimization module 403 is configured to optimize at least one feature interval according to the feature distribution characteristics of the multiple feature intervals to obtain an optimized second feature set.

[0115] The retrieval module 404 is configured to retrieve and classify the target query image based on the second feature set.

[0116] In a possible implementation, the interval division module 402 is configured to:

[0117] For each target category in the multiple categories, multiple image features contained in the target category are clustered to obtain a cluster center.

[0118] According to the distances from the multiple image features contained in the target category to the cluster center, the multiple image features are sorted to obtain a sorting result.

[0119] Based on the sorting results and the preset distance percentage step, multiple image features are divided into multiple feature intervals.

[0120] In a possible implementation, the optimization module 403 is configured to:

[0121] For each of the multiple feature intervals, a mean value of all image features in the feature interval is calculated to obtain representative features of the multiple feature intervals.

[0122] At least one representative feature is added to the first feature set to obtain an optimized second feature set.

[0123] In a possible implementation, the optimization module 403 is configured to:

[0124] The representative features of at least one feature interval among the multiple feature intervals are added separately to the first feature set to obtain an optimized second feature set. Or

[0125] Representative features of at least some of the feature intervals in the plurality of feature intervals are arranged and combined and added to the first feature set to obtain an optimized second feature set.

[0126] In a possible implementation, the optimization module 403 is configured to:

[0127] Count the number of image features in multiple feature intervals.

[0128] For a dense feature interval in which the number of image features exceeds a first threshold, the image features in the dense feature interval are compressed in quantity at a first compression rate to obtain an optimized second feature set.

[0129] In a possible implementation, the device further includes:

[0130] The compression module is configured to: for sparse feature intervals whose number is not greater than a second threshold, not perform image feature quantity compression or perform image feature quantity compression at a second compression rate to obtain an optimized second feature set, wherein the second compression rate is lower than the first compression rate.

[0131] In one possible implementation, the retrieval module 404 is configured to:

[0132] Determine query features from the target query image.

[0133] A similarity score is determined between the query feature and a plurality of features in the second feature set.

[0134] Determine the category to which the target query image belongs based on the preset voting strategy and similarity score.

[0135] In a possible implementation, the device further includes:

[0136] The voting strategy selection module is used to select the optimal voting strategy for multiple categories in the image database based on the historical retrieval performance data or feature distribution data of the category from multiple candidate voting strategies as the preset voting strategy.

[0137] In one possible implementation, the voting strategy selection module is used to:

[0138] The similarity scores are normalized using the Softmax function to obtain normalized scores.

[0139] The category with the highest normalized score is determined as the category to which the target query image belongs.

[0140] In one possible implementation, the voting strategy selection module is used to:

[0141] The similarity score is compared with a preset similarity threshold.

[0142] For each category, count the number of times the similarity score of its features exceeds the similarity threshold to get a vote.

[0143] The category with the highest number of votes is determined as the category to which the target query image belongs.

[0144] In one possible implementation, the voting strategy selection module is used to:

[0145] The similarity scores are sorted in descending order, and the top K similarity scores and their corresponding candidate categories are selected, where K is a preset integer greater than 1.

[0146] For each candidate category, the similarity scores of all the top K selected categories are accumulated to obtain a total category score.

[0147] The category with the highest total category score is determined as the category to which the target query image belongs.

[0148] The value of K is 3, 5 or 9.

[0149] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0150] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0151] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0152] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0153] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0154] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0155] The computing unit 501 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the optimization method for image search and image detection. For example, in some embodiments, the optimization method for image search and image detection can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the optimization method for image search and image detection described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the optimization method of image search and detection in any other appropriate manner (eg, by means of firmware).

[0156] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0157] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0158] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0159] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0160] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0161] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0162] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0163] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An image search optimization method, comprising: Acquire multiple image features of multiple categories from a first feature set corresponding to the image base library; Dividing the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space; Optimizing at least one feature interval according to the feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set; and Based on the second feature set, the target query image is retrieved and classified.

2. The method according to claim 1, wherein The dividing the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space comprises: For each target category in the multiple categories, clustering multiple image features included in the target category to obtain a cluster center; Sorting the plurality of image features according to the distances from the plurality of image features included in the target category to the cluster center to obtain a sorting result; According to the sorting result, the plurality of image features are divided into a plurality of ordered feature intervals.

3. The method according to claim 1, wherein The optimizing process for at least one feature interval according to the feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set includes: For each of the plurality of feature intervals, calculating a mean value of all image features within the feature interval to obtain representative features of the plurality of feature intervals; At least one of the representative features is added to the first feature set to obtain an optimized second feature set.

4. The method according to claim 3, wherein: Adding at least one representative feature to the first feature set to obtain an optimized second feature set includes: adding representative features of at least one feature interval among the plurality of feature intervals to the first feature set to obtain an optimized second feature set; or The representative features of at least some of the feature intervals in the plurality of feature intervals are arranged and combined and then added to the first feature set to obtain an optimized second feature set.

5. The method according to claim 1, wherein The optimizing process for at least one feature interval according to the feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set includes: Counting the number of image features within the plurality of feature intervals; For a dense feature interval in which the number of image features exceeds a first threshold, the image features in the dense feature interval are compressed in quantity at a first compression rate to obtain an optimized second feature set.

6. The method according to claim 5, further comprising: For sparse feature intervals whose number is not greater than a second threshold, the number of image features is not compressed or the image features are compressed at a second compression rate to obtain an optimized second feature set; wherein the second compression rate is less than the first compression rate.

7. The method according to claim 1, wherein The retrieving and classifying the target query image based on the second feature set includes: determining query features from a target query image; determining a similarity score between the query feature and a plurality of features in the second feature set; The category to which the target query image belongs is determined according to a preset voting strategy and the similarity score.

8. The method according to claim 1 or 7, further comprising: For multiple categories in the image database, based on historical retrieval performance data or feature distribution data of the categories, an optimal voting strategy is selected for the categories from multiple candidate voting strategies to serve as the preset voting strategy.

9. The method according to claim 7, wherein: The determining the category to which the target query image belongs according to a preset voting strategy and the similarity score includes: Normalizing the similarity score using a soft maximum function to obtain a normalized score; The category with the highest normalized score is determined as the category to which the target query image belongs.

10. The method according to claim 7, wherein: The determining the category to which the target query image belongs according to a preset voting strategy and the similarity score includes: comparing the similarity score with a preset similarity threshold; For each category, count the number of times the similarity score of its features exceeds the similarity threshold to get a vote; The category with the highest number of votes is determined as the category to which the target query image belongs.

11. The method according to claim 7, wherein: The determining the category to which the target query image belongs according to a preset voting strategy and the similarity score includes: Arrange the similarity scores in descending order and select the top K similarity scores and their corresponding candidate categories, where K is a preset integer greater than 1; For each candidate category, the similarity scores of the top K candidates are accumulated to obtain a total category score; Determine the category with the highest category total score as the category to which the target query image belongs; The value of K is 3, 5 or 9.

12. An image search optimization device, comprising: A feature acquisition module is used to acquire multiple image features of multiple categories from a first feature set corresponding to the image base library; An interval division module, configured to divide the plurality of image features into a plurality of feature intervals based on the distribution of the plurality of image features in the feature space; an optimization module, configured to optimize at least one feature interval according to feature distribution characteristics of the plurality of feature intervals to obtain an optimized second feature set; as well as A retrieval module is used to retrieve and classify the target query image based on the second feature set.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.