A method and device for selecting an image search model

By integrating the selection method of manual feature model and deep learning model, combining the parameter panel and online training mode, the problem of matching effect and computing efficiency in the graph search technology is solved, and efficient model selection and hardware resource optimization are achieved in changing business scenarios.

CN114332508BActive Publication Date: 2025-07-29BEIJING E HUALU INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210001497.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-04
Publication Date
2025-07-29
Estimated Expiration
2042-01-04

AI Technical Summary

Technical Problem

The existing picture-based search technology cannot take into account matching effects and computing efficiency, and a single model cannot meet the changing business needs and hardware resource limitations.

Method used

By integrating the selection method of manual feature models and deep learning models, combining parameter panels and online training modes, the optimal model is automatically selected and adapted to business changes, reducing hardware resource consumption.

Benefits of technology

It realizes optimization of taking into account matching effect and computing efficiency in changing business scenarios, reduces hardware resource requirements, and improves model adaptability and training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332508B_ABST
    Figure CN114332508B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for selecting an image search model. The method includes the following steps: retrieving a preset number of reference images with the highest similarity to each image to be retrieved from a reference image library using a handcrafted feature model; obtaining evaluation metrics of the handcrafted feature model; selecting the handcrafted feature model as the image search model when the evaluation metrics of the handcrafted feature model meet preset conditions; when the evaluation metrics of the handcrafted feature model do not meet the preset conditions, retrieving a preset number of reference images with the highest similarity to each image to be retrieved from the reference image library using a deep learning model; obtaining evaluation metrics of the deep learning model; and selecting the deep learning model as the image search model when the evaluation metrics of the deep learning model meet the preset conditions. Embodiments of the present application integrate the image matching process and the metric evaluation process, thereby taking into account both the matching effect and the computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a method and device for selecting an image search model. Background Art

[0002] With the development of the Internet and the digital economy, data is increasing at an exponential rate. As an important form of data, images widely exist in all walks of life in society, such as social networking, e-commerce, healthcare, transportation, and security. In these industries, the number of images is huge, and the task of querying similar images, i.e., searching for images by image, has attracted more and more attention.

[0003] Searching for images by image means searching for images with similar content based on the image content. The rapid growth of the data scale poses higher requirements for the calculation effect and calculation efficiency of searching for images by image. Building an image search system needs to solve two problems: the first is to extract image features, and the second is to build a database and provide a similarity search function. Currently, image search technologies usually use a single model or solution, and cannot balance the matching effect and calculation efficiency.

[0004] Application Content

[0005] The purpose of the embodiments of this application is to provide a method and device for selecting an image search model to solve the defect that the prior art cannot balance the matching effect and calculation efficiency.

[0006] To solve the above technical problems, this application is implemented as follows:

[0007] In a first aspect, a method for selecting an image search model is provided, including the following steps:

[0008] Set a reference image library and a set of images to be retrieved. The reference image library includes multiple reference images, the set of images to be retrieved includes multiple images to be retrieved, and there is a target image in the reference image library that matches each image to be retrieved;

[0009] Use a handcrafted feature model to retrieve a preset number of reference images with the highest similarity to each image to be retrieved from the reference image library;

[0010] Obtain the evaluation metrics of the handcrafted feature model according to the preset number of reference images;

[0011] When the evaluation metrics of the handcrafted feature model meet the preset conditions, select the handcrafted feature model as the image search model;

[0012] When the evaluation metrics of the manual feature model do not meet the preset conditions, use the deep learning model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library;

[0013] Obtain the evaluation metrics of the deep learning model according to the preset number of reference pictures;

[0014] When the evaluation metrics of the deep learning model meet the preset conditions, select the deep learning model as the picture search model.

[0015] In a second aspect, a device for selecting a picture search model is provided, including:

[0016] A setting module for setting a reference picture library and a set of pictures to be retrieved. The reference picture library includes multiple reference pictures, the set of pictures to be retrieved includes multiple pictures to be retrieved, and there is a target picture matching each picture to be retrieved in the reference picture library;

[0017] A first retrieval module for using the manual feature model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library;

[0018] A first acquisition module for obtaining the evaluation metrics of the manual feature model according to the preset number of reference pictures;

[0019] A first selection module for selecting the manual feature model as the picture search model when the evaluation metrics of the manual feature model meet the preset conditions;

[0020] A second retrieval module for using the deep learning model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library when the evaluation metrics of the manual feature model do not meet the preset conditions;

[0021] A second acquisition module for obtaining the evaluation metrics of the deep learning model according to the preset number of reference pictures;

[0022] A second selection module for selecting the deep learning model as the picture search model when the evaluation metrics of the deep learning model meet the preset conditions.

[0023] The embodiments of the present application integrate the picture matching process and the index evaluation process, so as to realize the reasonable selection of the manual feature model and the deep learning neural network model, and further balance the matching effect and the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of a method for selecting a picture search model provided by an embodiment of the present application;

[0025] Figure 2 It is a schematic diagram of retrieving reference pictures from the reference picture library using a manual feature model provided by an embodiment of the present application;

[0026] Figure 3 It is a schematic structural diagram of a deep learning model provided by an embodiment of the present application;

[0027] Figure 4 It is a schematic diagram of a parameter panel provided by an embodiment of the present application;

[0028] Figure 5 It is a schematic diagram of online generating picture data provided by an embodiment of the present application;

[0029] Figure 6 It is a schematic structural diagram of a selection device of a picture search model provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] For the scenario of searching pictures by pictures, the matching effect and calculation efficiency are equally important. For some simple scenarios, traditional non - deep learning models can already achieve good effects, and the calculation efficiency is far superior to that of deep learning models, and the requirements for hardware resources are also greatly reduced. The embodiments of the present application realize the search for the optimal model through the automation of evaluation, and reasonably select the manual feature model and the deep learning neural network model. At the same time, the picture matching process and the index evaluation process are integrated, and the picture calibration is exposed to the user in a visual way. The user only needs to provide data and click the mouse to quickly complete the model selection and testing.

[0032] In the business scenario of image search by image, the types of images to be matched often change over time. During actual use, new requirements will also arise. For example, in the cultural and entertainment social scenario, the influence of content such as subtitles and stickers needs to be avoided; in the field of traffic safety, the influence of image noise caused by seasons and sky colors needs to be avoided; in the field of security, the influence of image quality and image blurriness needs to be avoided, etc. This change in requirements is uncertain and occurs frequently. In fields related to public safety or business secrets, the model needs to be retrained and deployed in the local production environment. In the embodiments of the present application, various image change noises and disturbances are abstracted and classified through a parameter panel. Users make selections on the panel according to their own scenarios and determine the parameter thresholds for each image transformation. Through this visual operation logic, the code logic for generating training data is determined to adapt to the image distribution in the new actual business scenario, realizing the evolution and iteration of the model based on the parameter panel.

[0033] In addition, under normal circumstances, a full amount of training data needs to be prepared in advance when training a model. The data preparation and training processes are separated. For model training at the image search by image level, a large amount of data is often required to support it, bringing great pressure to the hard disk storage resources. In the embodiments of the present application, through the online training mode, only a limited amount of initial seed data needs to be provided. The generation of new data only exists in the process of generating training image batches, which are loaded in memory or video memory and are destroyed after updating the training weights. The online training method not only ensures the scale and diversity of training data but also ensures that the consumption of hard disk resources remains at the level of the initial seed data, realizing online training of the model based on a small amount of physical storage.

[0034] Next, in combination with the accompanying drawings, the method for selecting an image search model provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0035] As Figure 1 shown, it is a flowchart of a method for selecting an image search model provided by the embodiments of the present application. The method includes the following steps:

[0036] Step 101, set a reference image library and a set of images to be retrieved.

[0037] Among them, the reference image library includes multiple reference images, the set of images to be retrieved includes multiple images to be retrieved, and there is a target image in the reference image library that matches each image to be retrieved.

[0038] In this embodiment, two paths of picture folders need to be set, one is the path of the reference picture library, and the other is the path of the picture set to be retrieved; the purpose of testing pictures is to find the matching pictures corresponding to the pictures to be retrieved from the path of the reference picture library. The matching relationship can be given in a text format by the user in advance, as shown in Table 1; it can also be calibrated by the user in the subsequent visualization window.

[0039] File Name of Image 1 to be Retrieved \t File Name of Matched Image 1 in Reference Image Library File Name of Image 2 to be Retrieved \t File Name of Matched Image 2 in Reference Image Library ...... \t ...... File Name of Image n to be Retrieved \t File Name of Matched Image n in Reference Image Library

[0040] Table 1 Matching relationship in text form

[0041] Step 102, use the handcrafted feature model to retrieve a preset number of reference pictures with the highest similarity from the reference picture library for each picture to be retrieved.

[0042] Specifically, the handcrafted feature model can be used to generate the feature codes of each reference picture in the reference picture library and each picture to be retrieved in the picture set to be retrieved; calculate the Euclidean distance between the feature code of each picture to be retrieved and the feature code of each reference picture in the reference picture library; according to the Euclidean distance, determine a preset number of reference pictures with the highest similarity to each picture to be retrieved.

[0043] Among them, the handcrafted feature model includes two types of algorithms: the image hashing algorithm and the GIST algorithm. Each algorithm can generate a fixed-length feature code and is used to calculate the Euclidean distance. Finally, the similarity between two pictures can be obtained according to the distance between the feature codes of the two pictures.

[0044] For the image hashing algorithm, the specific steps include:

[0045] (1) Resize the picture: Resize the picture to 9*8 size, with a total of 72 pixels.

[0046] (2) Convert to grayscale: Convert the picture into a single-channel grayscale picture.

[0047] (3) Calculate the difference value: Calculate the difference matrix G, and each element of G is the pixel value of the current row - the pixel value of the previous row. The size of G is 8*8.

[0048] (4) Calculate the picture code: Initialize the dhash of the input picture as "", and traverse each pixel of the matrix G from left to right and top to bottom. When the element G(i, j) in the i-th row and j-th column >= a, then dhash += "1"; when the element G(i, j) in the i-th row and j-th column < a, then dhash += "0".

[0049] (5) The similarity score of two pictures is defined as: similarity = (64 - Hamming distance between two pictures) / 64.

[0050] For the GIST algorithm (512 dimensions), the specific steps are as follows:

[0051] (1) Use 32 Gabor filters to perform convolution at 4 scales and 8 directions to obtain 32 feature maps, and the size of the feature maps is the same as that of the input image;

[0052] (2) Divide each feature map into 16 regions of 4*4, and calculate the mean value of each region;

[0053] (3) 32 feature maps, each containing 16 mean values, form a 512-dimensional GIST feature encoding;

[0054] (4) The similarity score of two pictures is defined as: similarity = 1 / (1 + Euclidean distance between the two pictures).

[0055] Step 103, obtain the evaluation metrics of the manual feature model according to the preset number of reference pictures.

[0056] Among them, the evaluation metrics include hit rate and score; correspondingly, it is possible to determine the hit rate of the manual feature model according to whether the target picture corresponding to the to-be-retrieved picture is included in the preset number of reference pictures with the highest similarity to each to-be-retrieved picture; in the case where the target picture corresponding to the to-be-retrieved picture is included in the preset number of reference pictures with the highest similarity to each to-be-retrieved picture, determine the score of the manual feature model according to the similarity ranking of the target picture in the preset number of reference pictures.

[0057] In this embodiment, the index evaluation is used to determine whether the model meets the user's requirements. The user provides a reference picture library and a set of to-be-retrieved pictures. The system will automatically calculate the feature encoding of each picture and give the top K most matching pictures and scores of each to-be-retrieved picture, as Figure 2 shown. When the user provides a text file of the matching relationship, the system will automatically calculate each index; when the user does not provide it, the system will visually give each to-be-retrieved picture and the top K matching pictures calculated by the system. For each to-be-retrieved picture, the user calibrates the K candidate pictures given by the system. When the k-th picture is hit, calibrate the picture. When none of the K pictures is hit, calibrate "none". The value of K can be set by the user himself, and the system defaults to K = 5.

[0058] Among them, the evaluation metrics include two items: (1) top K hit rate; (2) score. The calculation method of the top K hit rate is: top K hit rate = number of to-be-retrieved pictures hit / total number of to-be-retrieved pictures; the score calculation method is: Among them, Rank represents the descending order value of the matching pictures that are hit among the K matching pictures. For each manual feature model and each trained deep learning model, steps 101 to 103 will be executed one to multiple times until a model that meets the user's conditions is obtained.

[0059] Step 104, when the evaluation metrics of the manual feature model meet the preset conditions, select the manual feature model as the image search model.

[0060] Step 105, when the evaluation metrics of the manual feature model do not meet the preset conditions, use the deep learning model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library.

[0061] Specifically, the deep learning model uses the Residual Network Res50 to obtain the feature map, uses the GeM pooling layer to increase the contrast of the feature map, making it pay more attention to the saliency map. After passing through the normalization layer, the feature encoding is trained using the triplet loss function to update the model, as Figure 3 shown. Among them, the structural parameters of the ResNet50 feature extraction backbone are shown in Table 2.

[0062]

[0063] Table 2 Structural parameters of the ResNet50 feature extraction backbone

[0064] The GeM layer is defined as:

[0065]

[0066] Among them, is the feature map calculated by Res50, C is the number of channels, H is the height, and W is the width; u ∈ Ω = {1,..., H} × {1,..., W}, which can be regarded as a pixel of the feature map; p is a parameter, and its value range is between 1 and positive infinity. In this system, P takes the value of 3.

[0067] The triplet loss function is defined as:

[0068] L(e i , e j , β, y ij ) = max{0, α + y ij (D(e i , e j ) - β)}

[0069] Among them, i and j are picture codes, ei and ej represent the feature encodings corresponding to the pictures, α and β are hyperparameters to be learned, y ijrepresents the matching relationship between Picture i and Picture j, with a match being 1 and a non-match being -1. D(e i , e j ) is the Euclidean distance after normalizing the feature encodings of the two pictures:

[0070]

[0071] The steps for the deep learning model to calculate the similarity of pictures are as follows:

[0072] (1) Convert the picture size to an RGB three-channel image and normalize it by channel; the normalization formula is:

[0073]

[0074] where μ = [0.485, 0.456, 0.406] and std = [0.229, 0.224, 0.225].

[0075] (2) Input into the deep learning model to obtain the feature encoding (1500 dimensions) of the picture;

[0076] (3) The similarity score of two pictures is defined as: Similarity = 1 / (1 + the Euclidean distance between the two pictures).

[0077] Step 106: Obtain the evaluation metrics of the deep learning model according to the preset number of reference pictures.

[0078] Step 107: Select the deep learning model as the picture search model when the evaluation metrics of the deep learning model meet the preset conditions.

[0079] The embodiment of this application integrates the picture matching process and the metric evaluation process, so as to realize the reasonable selection of the manual feature model and the deep learning neural network model, and further balance the matching effect and calculation efficiency.

[0080] Furthermore, in the embodiment of this application, after step 106, the following steps can also be executed:

[0081] Step 108: When the evaluation metrics of the deep learning model do not meet the preset conditions, update the training picture set according to the picture transformation type selected by the user in the parameter panel.

[0082] Among them, the picture transformation types include layer covering, spatial transformation, pixel jitter, and color transformation. Layer covering includes text covering and picture covering; spatial transformation includes cropping, rotation, horizontal flipping, vertical flipping, padding, aspect ratio, and perspective transformation; pixel jitter includes blurring, encoding compression, sharpening, mosaicking, and pixel washing; color transformation includes brightness, saturation, contrast, and grayscale.

[0083] Step 109: Based on the updated training image set, perform online training on the manual feature model and the deep learning model.

[0084] It should be noted that after the online training of the manual feature model and the deep learning model is completed, steps 101 and its subsequent steps can be returned to perform to select the manual feature model and the deep learning model.

[0085] In the embodiments of the present application, when the pre-loaded deep learning model cannot meet the user's needs, it is often because the training data of the pre-loaded deep learning model is inconsistent with the distribution of the image data in the actual application scenario. In this case, new image data needs to be used to train the model.

[0086] In addition, during the actual business use, the image distribution will shift or fluctuate, and in addition, there will also be new matching type requirements. In this case, new image data needs to be used to fine-tune the model;

[0087] In the above two cases, since the customer data may involve factors such as business secrets, it is necessary to perform re-training and deployment in the original environment. In the embodiments of the present application, in the form of a parameter panel, noise and perturbations are classified and quantified, allowing users to directly define the required image transformation types and generate a new image transformation mode with one key, such as Figure 4 as shown.

[0088] The embodiments of the present application incorporate common image transformation methods into the parameter panel, mainly including four transformation methods: layer covering, spatial transformation, pixel jittering, and color transformation. Under layer covering, there are two operation forms, namely text covering and image covering; under spatial transformation, there are seven operation methods, namely cropping, rotation, horizontal flipping, vertical flipping, padding, aspect ratio, and perspective transformation; under pixel jittering, there are five operation forms, namely blurring, encoding compression, sharpening, mosaicking, and pixel scrubbing; under color transformation, there are four operation methods, namely brightness, saturation, contrast, and grayscale. For each operation method, a probability value can be set, which is between 0 and 1; for operation methods with a parameter range, the change range of the parameters can also be set on the parameter panel, usually in the form of a maximum value and a minimum value. When actually generating data, the selected parameter is a random value that changes between the maximum value and the minimum value.

[0089] After the transformation method is determined, the code logic for generating the training image data will be determined, and no new image data will be generated. After the parameters are determined through the parameter panel, the model is trained and updated. In general, the training data needs to be stored in full on the hard disk. For services such as image search, the amount of data is usually more than a few hundred GB. When training is required at the project site, the requirements for hard disk resources are very high. The embodiment of the present application uses the method of online generation of image data for training, such as Figure 5 As shown, newly generated training data is not stored on disk, but only in memory or video memory for training. The system is pre-installed with an initial set of 10,000 seed images, which are stored on disk. During training, a batch of training images is divided into two parts: "single images" and "remaining image data." "Single images" regenerate a batch of transformed data based on the user's settings on the parameter panel. These are represented by codes like "1_1," "1_2," "1_3," and "1_4," representing differently transformed copies of the same image. "Remaining image data" undergoes only simple data augmentation and is represented by codes like "2," "3," "4," and "5." These images together form a new batch of images, which is stored in memory or video memory. "1_1," "1_2," "1_3," and "1_4" have a matching relationship, but they do not match "2," "3," "4," and "5." After the model weights are updated using the loss function, the online generated image batch is immediately destroyed.

[0090] This embodiment of the application follows the principle of Occam's razor, using the most efficient model possible while ensuring model effectiveness. This is particularly important for image search systems with huge data volumes. By integrating indicator calculations and visual calibration, the model selection process is automated, reducing user learning costs.

[0091] In addition, after the image search system is actually launched, there will usually be overall image distribution deviation and disturbance, and the project site itself will also have new image matching requirements. This embodiment of the application models and quantifies common noise disturbance methods and opens it to users in the form of a parameter panel. Users can set parameters on the parameter panel according to their actual situation. The system automatically generates corresponding training data code logic to achieve noise and disturbance adaptation.

[0092] Furthermore, considering that image-based image search training data typically takes up a significant amount of hard drive storage space, the present embodiment utilizes an online training model based on a small amount of hard drive storage. This model only requires a small amount of raw training data, and then uses parameters determined by a parameter panel to augment new training data online into the memory or video memory. The data is then destroyed after the model parameters are updated. This training approach not only reduces the use of hard drive storage space but also ensures the scale and diversity of the data during training.

[0093] As shown in Figure 6 the figure, it is a schematic structural diagram of a selection device for an image search model provided by an embodiment of the present application, including:

[0094] A setting module 610, configured to set a reference image library and a set of images to be retrieved. The reference image library includes multiple reference images, the set of images to be retrieved includes multiple images to be retrieved, and there is a target image in the reference image library that matches each image to be retrieved;

[0095] A first retrieval module 620, configured to retrieve a preset number of reference images with the highest similarity to each image to be retrieved from the reference image library using a handcrafted feature model;

[0096] Specifically, the first retrieval module 620 is specifically configured to: generate feature encodings of each reference image in the reference image library and each image to be retrieved in the set of images to be retrieved using a handcrafted feature model; calculate the Euclidean distance between the feature encoding of each image to be retrieved and the feature encoding of each reference image in the reference image library; and determine a preset number of reference images with the highest similarity to each image to be retrieved according to the Euclidean distance.

[0097] A first acquisition module 630, configured to obtain an evaluation metric of the handcrafted feature model according to the preset number of reference images;

[0098] Wherein, the evaluation metrics include hit rate and score.

[0099] Specifically, the first acquisition module 630 is specifically configured to: determine the hit rate of the handcrafted feature model according to whether the preset number of reference images with the highest similarity to each image to be retrieved includes the target image corresponding to the image to be retrieved; and determine the score of the handcrafted feature model according to the similarity ranking of the target image in the preset number of reference images in the case where the preset number of reference images with the highest similarity to each image to be retrieved includes the target image corresponding to the image to be retrieved.

[0100] A first selection module 640, configured to select the handcrafted feature model as the image search model when the evaluation metrics of the handcrafted feature model meet a preset condition;

[0101] A second retrieval module 650, configured to retrieve a preset number of reference images with the highest similarity to each image to be retrieved from the reference image library using a deep learning model when the evaluation metrics of the handcrafted feature model do not meet a preset condition;

[0102] A second acquisition module 660, configured to obtain an evaluation metric of the deep learning model according to the preset number of reference images;

[0103] A second selection module 670, configured to select the deep learning model as an image search model when evaluation metrics of the deep learning model meet a preset condition.

[0104] Further, the above-mentioned apparatus further includes:

[0105] An update module, configured to update a training image set according to an image transformation type selected by a user in a parameter panel when the evaluation metrics of the deep learning model do not meet the preset condition; wherein the image transformation type includes layer covering, spatial transformation, pixel jittering, and color transformation;

[0106] A training module, configured to perform online training on the manual feature model and the deep learning model based on the updated training image set.

[0107] Wherein, the layer covering includes text covering and image covering; the spatial transformation includes cropping, rotation, horizontal flipping, vertical flipping, padding, aspect ratio, and perspective transformation; the pixel jittering includes blurring, encoding compression, sharpening, mosaicking, and pixel scrubbing; the color transformation includes brightness, saturation, contrast, and grayscale.

[0108] Embodiments of the present application integrate an image matching process and an index evaluation process, so as to realize a reasonable selection of a manual feature model and a deep learning neural network model, and further balance matching effects and computing efficiency.

[0109] Embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned method embodiment for selecting an image search model is implemented, and the same technical effects can be achieved. To avoid repetition, details are not described herein again. Wherein, the computer-readable storage medium, such as a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, or an optical disc, etc.

[0110] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0111] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0112] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A method for selecting a picture search model, characterized in that Including the following steps: Set up a reference picture library and a picture set to be retrieved. The reference picture library includes multiple reference pictures, the picture set to be retrieved includes multiple pictures to be retrieved, and there is a target picture in the reference picture library that matches each picture to be retrieved; Use a handcrafted feature model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library; Obtain the evaluation metrics of the handcrafted feature model based on the preset number of reference pictures; When the evaluation metrics of the handcrafted feature model meet the preset conditions, select the handcrafted feature model as the picture search model; When the evaluation metrics of the handcrafted feature model do not meet the preset conditions, use a deep learning model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library; Obtain the evaluation metrics of the deep learning model based on the preset number of reference pictures; When the evaluation metrics of the deep learning model meet the preset conditions, select the deep learning model as the picture search model; The step of using a handcrafted feature model to retrieve a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library specifically includes: Use a handcrafted feature model to generate feature encodings of each reference picture in the reference picture library and each picture to be retrieved in the picture set to be retrieved; Calculate the Euclidean distance between the feature encoding of each picture to be retrieved and the feature encodings of each reference picture in the reference picture library; Determine a preset number of reference pictures with the highest similarity to each picture to be retrieved based on the Euclidean distance.

2. The method according to claim 1, wherein The evaluation metrics include hit rate and score; The step of obtaining the evaluation metrics of the handcrafted feature model based on the preset number of reference pictures specifically includes: Determine the hit rate of the handcrafted feature model based on whether the preset number of reference pictures with the highest similarity to each picture to be retrieved includes the target picture corresponding to the picture to be retrieved; When the preset number of reference pictures with the highest similarity to each picture to be retrieved includes the target picture corresponding to the picture to be retrieved, determine the score of the handcrafted feature model according to the similarity ranking of the target picture in the preset number of reference pictures.

3. The method according to claim 1, wherein After obtaining the evaluation metrics of the deep learning model based on the preset number of reference pictures, it further includes: When the evaluation metrics of the deep learning model do not meet the preset conditions, update the training picture set according to the picture transformation type selected by the user in the parameter panel; where the picture transformation type includes layer covering, spatial transformation, pixel jitter, and color transformation; Based on the updated training picture set, perform online training on the handcrafted feature model and the deep learning model.

4. The method according to claim 3, characterized in that, The layer covering includes text covering and picture covering; the spatial transformation includes cropping, rotation, horizontal flipping, vertical flipping, padding, aspect ratio, and perspective transformation; the pixel jitter includes blurring, encoding compression, sharpening, mosaicking, and pixel scrubbing; the color transformation includes brightness, saturation, contrast, and grayscale.

5. A selection device for a picture search model, characterized in that, Including: A setting module for setting a reference picture library and a picture set to be retrieved. The reference picture library includes multiple reference pictures, and the picture set to be retrieved includes multiple pictures to be retrieved. Moreover, there is a target picture in the reference picture library that matches each picture to be retrieved. A first retrieval module for retrieving a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library using a handcrafted feature model. A first acquisition module for obtaining evaluation metrics of the handcrafted feature model based on the preset number of reference pictures. A first selection module for selecting the handcrafted feature model as a picture search model when the evaluation metrics of the handcrafted feature model meet preset conditions. A second retrieval module for retrieving a preset number of reference pictures with the highest similarity to each picture to be retrieved from the reference picture library using a deep learning model when the evaluation metrics of the handcrafted feature model do not meet preset conditions. A second acquisition module for obtaining evaluation metrics of the deep learning model based on the preset number of reference pictures. A second selection module for selecting the deep learning model as a picture search model when the evaluation metrics of the deep learning model meet preset conditions. The first retrieval module is specifically configured to: generate feature encodings of each reference picture in the reference picture library and each picture to be retrieved in the picture set to be retrieved using a handcrafted feature model. Calculate the Euclidean distance between the feature encoding of each picture to be retrieved and the feature encodings of each reference picture in the reference picture library; and determine a preset number of reference pictures with the highest similarity to each picture to be retrieved based on the Euclidean distance.

6. The device according to claim 5, characterized in that The evaluation metrics include hit rate and score. The first acquisition module is specifically configured to: determine the hit rate of the handcrafted feature model based on whether the preset number of reference pictures with the highest similarity to each picture to be retrieved includes the target picture corresponding to the picture to be retrieved. When the preset number of reference pictures with the highest similarity to each picture to be retrieved includes the target picture corresponding to the picture to be retrieved, determine the score of the handcrafted feature model according to the similarity ranking of the target picture in the preset number of reference pictures.

7. The device according to claim 5, characterized in that, It further includes: An update module for updating the training picture set according to the picture transformation type selected by the user in the parameter panel when the evaluation metrics of the deep learning model do not meet preset conditions; wherein the picture transformation types include layer overlay, spatial transformation, pixel jitter, and color transformation. A training module for performing online training on the handcrafted feature model and the deep learning model based on the updated training picture set.

8. The device according to claim 7, characterized in that The layer overlay includes text overlay and picture overlay; the spatial transformation includes cropping, rotation, horizontal flipping, vertical flipping, padding, aspect ratio, and perspective transformation; the pixel jitter includes blurring, encoding compression, sharpening, mosaicking, and pixel scrubbing; the color transformation includes brightness, saturation, contrast, and grayscale.

Citation Information

Patent Citations

  • Public security detection application-oriented image retrieval method

    CN108363771A

  • Method for artificial intelligence (AI) model selection

    WO2021195689A1